Antigen-antibody interaction prediction remains a fundamental challenge in computational biology and AI-driven drug discovery. While recent machine learning approaches have enabled accurate single-chain protein structure prediction, modeling antigen-antibody interactions remains challenging because of extreme sequence diversity, evolutionary asymmetry, conformational flexibility, and, critically, the lack of large experimental datasets with functional labels and high-resolution structural supervision tailored to antigen-antibody interactions. As a result, existing models often fail to generalize beyond narrow benchmarks and curated structural complexes. We announce SEPIQ (Sequence-based Epitope Prediction Intelligent Query), a community challenge built around a large experimental dataset of immune-derived VHH nanobody sequences. SEPIQ comprises two complementary tracks: Track A, managed by SEPIQ, where participants submit per-antigen-residue epitope/paratope probabilities, and Track B, managed by CAPRI (Critical Assessment of PRediction of Interactions), where participants submit predicted antigen-VHH complex structures according to CAPRI procedures. A key strength of SEPIQ is the availability of high-resolution Cryo-EM ground truth for blind evaluation with a large-scale antigen-VHH interaction dataset. To ensure rigor, Cryo-EM structures and derived test labels will be held out during the challenge and released only after completion of the submission evaluation. The SEPIQ challenge aims to catalyze “Sequence-to-Structure Immunology” by testing a “folding from binding” hypothesis: large-scale binding-derived sequence data may contain statistical signatures of structural and functional constraints when combined with limited high-resolution structural ground truth.
Accurate epitope identification and prediction of antigen-antibody interactions are critical for drug discovery, as they enable therapeutic design, define mechanisms of action, and allow assessment of efficacy and potential off-target effects. Despite their importance, reliable prediction of antigen-antibody interactions remains a central and unresolved challenge in computational biology. Recent advancements in machine learning led to the development of methods that excel at single-chain protein structure prediction, yet they remain limited for antigen-antibody modeling due to repertoire-scale sequence diversity, evolutionary asymmetry at interfaces, limited cross-chain coevolutionary signal in MSAs, epitope conformational variability, and binder/non-binder discrimination across large experimental panels. As a result, state-of-the-art models (e.g., AlphaFold3, OpenFold3, Boltz-1, and Chai-1) achieve acceptable accuracy on only a minority of cases in antigen-antibody docking benchmarks (40%), in contrast to their much higher success rates on single-chain protein structure prediction (≈90%) [The OpenFold3 Team, 2025, Yin and Pierce, 2024, Gaudreault et al., 2025, Hitawala and Gray, 2025, Abramson et al., 2024, Discovery et al., 2024, Jumper et al., 2021]. One of the key limiting factors in improving model performance for antigen-antibody complexes is the scarcity of experimental data that provides sequence diversity, functional binding outcomes, and high-resolution structural ground truth. The two largest publicly available datasets of antigen-antibody complex structures, SAbDab and SAAINT-DB, are curated from entries in the Protein Data Bank (PDB) [Huang et al., 2025, Schneider et al., 2022, Dunbar et al., 2014]. According to SAbDab, as of July 2026, the PDB contained 9,775 antibody structures with antigens, which underscores the limited scale of structural supervision available for predicting antigen-antibody complexes. Antigen-VHH complexes are even rarer (SAbDab-nano reports 2,013 structures with antigen). To accelerate the progress in prediction of antigen-antibody interactions, we announce the open SEPIQ machine learning challenge and provide novel large-scale experimental sequence data to enable robust modeling of antigen-antibody complexes, together with high-quality structural ground truth for evaluation, with test labels held out for blind assessment.
The competition will feature two primary tracks at different resolutions: Track A (residue-level resolution; SEPIQ) asks participants to develop sequence-based models that predict epitope/paratope residues (as per-residue probabilities or binary labels) from VHH and antigen sequence inputs, enabling interface characterization even without explicit 3D complex modeling. Track B (atomic-level resolution; CAPRI) asks participants to predict antigen-antibody complex structures from VHH and antigen sequences. We believe that the SEPIQ challenge will encourage the development of novel sequence-based architectures and has the potential to establish a new research subfield at the intersection of machine learning, immunology, and structural biology, accelerating the discovery of new medicines. Based on the social media outreach and positive feedback from in-person presentations, we expect more than 100 participants from diverse backgrounds to take part in the SEPIQ challenge per a year. Conceptually, SEPIQ tests a “folding from binding” hypothesis: large-scale sequence-level binding data may provide weak but scalable information about structural and functional constraints at antigen-antibody interfaces. We present this as a testable hypothesis rather than a definitive claim, because high-resolution structural labels remain limited and are used only for blind evaluation.
This challenge involves antigen-antibody complexes containing novel epitopes and paratopes. Accurate prediction of epitope-paratope interactions can accelerate therapeutic nanobody discovery by reducing the need for costly wet-lab screening and structural characterization. Sequence-only methods are particularly valuable in settings where experimentally determined structures are unavailable.
Track A (residue-level resolution) asks participants to predict which residues on the antigen (epitope) and the antibody (paratope) form the binding interface of each antigen-VHH complex, using sequence input only. Compared to three-dimensional structural data, sequence-based data on antigen-antibody interactions can be acquired more extensively and at a lower cost. Such data will continue to accumulate at an accelerating rate. The primary objective of Track A is to explore the potential applications of these substantial sequence-based data on antigen-antibody interactions. However, there are no restrictions on the algorithms used to predict epitopes/paratopes for Track A. It is perfectly acceptable to predict epitopes/paratopes based on three-dimensional structure prediction (equivalent to Track B). The primary submission format is a per-residue epitope/paratope probability vector. Please refer to the section below for a more detailed explanation. AI models developed for Track A would be able to identify antigen-antibody interfaces directly from sequence data, enabling antibody design without experimental structural information. We demonstrate the possibility of solving the challenge with our lightweight baseline model, which is trained on a publicly available dataset, in the 'Baseline and code' section.
Track B (atomic-level resolution) asks participants to predict antigen-VHH complex structures from antigen and VHH sequences. This track will follow the established CAPRI framework for target release, model submission, and assessment. In this track, participants should submit predicted complex atomic structures. Please wait for a future announcement from CAPRI for a more detailed explanation. Although several existing models, including Boltz-1, Chai-1, and AlphaFold3, can predict antigen-antibody complex structures, their accuracy on this task remains limited. Improving these models would advance computational modeling of antigen-antibody interactions and accelerate drug development.
Although the tasks are highly challenging due to sparse experimental supervision and the complexity of immune recognition, recent advances in protein foundation models, geometric deep learning, and antibody structure prediction demonstrate that meaningful progress is achievable. The aim of this challenge is to make antibody drug discovery more efficient in the real world. As long as the answer is correct, we do not consider data leakage resulting from the use of publicly available data to be an issue.
For each complex sequence, provide a table with chain names (chain A for the antigen, chain B for the VHH), residue sequence numbers, and amino acid residue names.
Participants must return tables formed by adding a column to each input table in csv format. The added column must contain values from 0 to 1 indicating the probability that the residue is part of the binding interface.
Example of output for the antigen:
chain | seqNum | resName | prediction |
A | 1 | ALA | 0 |
A | 2 | ILE | 0 |
A | 3 | ASN | 0 |
A | 4 | ALA | 0.95 |
A | 5 | ALA | 0.91 |
A | 6 | ARG | 0 |
A | 7 | GLU | 0 |
A | 8 | ASN | 0 |
A | 9 | LEU | 0.59 |
A | 10 | GLY | 0.15 |
A | 11 | ALA | 0.17 |
A | 12 | SER | 0 |
B | 1 | MET | 0 |
B | 2 | ALA | 0 |
B | 3 | GLN | 0.87 |
B | 4 | VAL | 0.21 |
B | 5 | GLN | 0.74 |
B | 6 | LEU | 0 |
B | 7 | GLN | 0 |
B | 8 | GLU | 0.47 |
B | 9 | SER | 0.14 |
B | 10 | GLY | 0 |
B | 11 | GLY | 0 |
B | 12 | GLY | 0 |
Participants may skip residues in their output — it will be assumed that the prediction is 0 for any omitted residue.
Example of shorter output for the antigen:
chain | seqNum | resName | prediction |
A | 4 | ALA | 0.95 |
A | 5 | ALA | 0.91 |
A | 9 | LEU | 0.59 |
A | 10 | GLY | 0.15 |
A | 11 | ALA | 0.17 |
B | 3 | GLN | 0.87 |
B | 4 | VAL | 0.21 |
B | 5 | GLN | 0.74 |
B | 8 | GLU | 0.47 |
B | 9 | SER | 0.14 |
The template CSV files for predictions have been uploaded to the Hugging Face repository.
Given an experimentally determined antigen-VHH complex structure, the interface will be described using the Voronota-LT method, which constructs Voronoi tessellation-based inter-atom contact surfaces and calculates their areas.
We will adopt the Voronota-LT software [Kliment et al., 2025] to calculate only the relevant contacts — that is, the contacts between the antigen and the VHH chains. Moreover, we can instruct Voronota-LT to summarize (aggregate) contact areas for each residue that has at least one relevant contact.
Let us illustrate the interface determination process using a known antigen-VHH structure from the Protein Data Bank (PDB).
Firstly, let us obtain the software:
# download the universal Voronota-LT version 1.1.479 executable
wget "https://github.com/kliment-olechnovic/voronota/releases/download/v1.29.4602/cosmopolitan_voronota-lt_v1.1.479.exe"
# rename the downloaded executable
mv cosmopolitan_voronota-lt_v1.1.479.exe voronota-lt
# set execute permission
chmod +x ./voronota-ltSecondly, let us prepare an input structure — download and unzip the first assembly structure from the PDB entry 5F7L (https://www.rcsb.org/structure/5F7L):
wget "https://files.rcsb.org/download/5F7L-assembly1.cif.gz"
gunzip ./5F7L-assembly1.cif.gzWe note that in the downloaded structure there are two chains: chain A is the antigen, and chain B is the VHH.
Finally, let us run Voronota-LT to produce a table of residue-level binding site areas:
./voronota-lt --input "./5F7L-assembly1.cif" \
--restrict-contacts '[-a1 [-chain A] -a2 [-chain B]]' \
--write-sites-residue-level-to-file ./sites-residue-level.tsvAfter that, the generated file sites-residue-level.tsv contains the binding site areas for all residues that participate in the interface:
rs_header | ID_chain | ID_rnum | ID_rname | area | arc_length | distance | count |
rs | A | 102 | TYR | 26.2887 | 5.35143 | 3.25509 | 10 |
rs | A | 103 | ALA | 5.81737 | 2.34553 | 3.9238 | 3 |
rs | A | 104 | THR | 10.2041 | 1.92648 | 3.50697 | 3 |
rs | A | 105 | GLN | 52.2302 | 6.67633 | 3.06804 | 31 |
rs | A | 106 | CYS | 12.9647 | 0 | 2.69348 | 6 |
rs | A | 116 | THR | 19.9032 | 6.73551 | 4.12098 | 13 |
rs | A | 117 | SER | 12.8987 | 15.7238 | 4.87017 | 9 |
rs | A | 119 | THR | 22.7757 | 11.912 | 3.51999 | 10 |
rs | A | 121 | ILE | 15.189 | 7.36034 | 4.00473 | 3 |
rs | A | 128 | TYR | 66.5253 | 20.6576 | 3.23458 | 34 |
rs | A | 134 | THR | 50.6727 | 14.9325 | 2.83448 | 24 |
rs | A | 135 | CYS | 2.61009 | 0 | 3.58477 | 3 |
rs | A | 136 | SER | 39.7118 | 3.33127 | 2.52272 | 25 |
rs | A | 137 | LEU | 53.1561 | 0 | 3.44333 | 35 |
rs | A | 138 | ASN | 42.8299 | 5.52968 | 2.7755 | 25 |
rs | A | 139 | ARG | 81.9165 | 19.6108 | 3.05166 | 43 |
rs | A | 140 | TYR | 86.4117 | 17.5458 | 2.50269 | 39 |
rs | A | 145 | TYR | 9.02516 | 5.71504 | 4.12956 | 5 |
rs | A | 146 | GLY | 8.25783 | 2.21431 | 4.67829 | 4 |
rs | A | 147 | PRO | 8.5554 | 0.783521 | 4.72983 | 5 |
rs | A | 190 | SER | 9.85719 | 4.90913 | 4.31282 | 4 |
rs | A | 191 | GLY | 15.6605 | 1.62435 | 4.07761 | 5 |
rs | A | 192 | GLU | 2.93553 | 4.35852 | 5.05939 | 2 |
rs | A | 241 | THR | 15.8578 | 6.29968 | 4.28125 | 7 |
rs | A | 242 | ARG | 38.0833 | 24.4135 | 4.69196 | 14 |
rs | A | 243 | VAL | 32.7166 | 18.5909 | 3.30051 | 9 |
rs | A | 279 | TYR | 46.6128 | 3.07466 | 2.54624 | 28 |
rs | A | 281 | SER | 43.2088 | 8.06525 | 2.49241 | 15 |
rs | A | 282 | VAL | 11.9861 | 3.03489 | 3.59902 | 9 |
rs | A | 283 | THR | 30.5658 | 7.36819 | 3.50735 | 16 |
rs | A | 295 | ARG | 54.2229 | 21.4632 | 3.40346 | 24 |
rs | B | 1 | GLN | 97.4462 | 18.6789 | 2.66855 | 57 |
rs | B | 2 | VAL | 13.6402 | 2.62611 | 2.52272 | 8 |
rs | B | 3 | GLN | 56.1574 | 21.8633 | 3.51999 | 19 |
rs | B | 5 | GLN | 50.7106 | 26.0035 | 4.28125 | 22 |
rs | B | 6 | GLU | 0.853008 | 1.11928 | 5.77592 | 1 |
rs | B | 7 | SER | 8.57829 | 8.99021 | 5.43684 | 1 |
rs | B | 26 | GLY | 24.476 | 11.6387 | 3.23458 | 7 |
rs | B | 27 | SER | 23.1216 | 3.52698 | 3.94957 | 11 |
rs | B | 28 | ILE | 2.73042 | 1.68464 | 4.77136 | 2 |
rs | B | 29 | TYR | 99.7546 | 17.6416 | 2.49241 | 54 |
rs | B | 30 | SER | 13.7472 | 6.48282 | 3.24403 | 5 |
rs | B | 45 | HIS | 17.0975 | 8.7078 | 4.61964 | 9 |
rs | B | 99 | SER | 14.3742 | 0 | 3.19291 | 4 |
rs | B | 100 | ASP | 49.1928 | 8.82601 | 2.54624 | 23 |
rs | B | 101 | ARG | 109.806 | 20.2009 | 2.97931 | 65 |
rs | B | 102 | LEU | 13.9772 | 0.810522 | 3.74136 | 7 |
rs | B | 103 | THR | 20.9277 | 7.30925 | 3.96832 | 9 |
rs | B | 104 | ASP | 6.27695 | 6.1968 | 4.84443 | 2 |
rs | B | 108 | CYS | 16.5484 | 6.77437 | 3.23494 | 5 |
rs | B | 109 | GLU | 50.0274 | 12.2063 | 2.50269 | 19 |
rs | B | 110 | ALA | 7.51728 | 0 | 3.85594 | 5 |
rs | B | 111 | ASP | 47.5298 | 0 | 2.7755 | 28 |
rs | B | 112 | TYR | 43.1371 | 1.95946 | 3.34279 | 27 |
rs | B | 113 | TRP | 71.6285 | 19.3995 | 3.60848 | 46 |
rs | B | 114 | GLY | 3.13616 | 5.68057 | 5.47438 | 3 |
rs | B | 115 | GLN | 67.2589 | 33.2266 | 3.30051 | 24 |
From that table, we extract only the required columns (ID_chain, ID_rnum, ID_rname, area), rename them to (chain, seqNum, resName, weight) respectively, insert an additional fact column with all values set to 1 (because all reported residues are interface residues), and generate a table corresponding to the ground truth format:
{
echo "chain seqNum resName fact weight" | sed 's/\s\+/\t/g'
cat "./sites-residue-level.tsv" | tail -n +2 | awk '{print $2 "\t" $3 "\t" $4 "\t1\t" $5}'
} \
> "./positive_ground_truth.tsv"This will produce a table of positive ground truth:
chain | seqNum | resName | fact | weight |
A | 102 | TYR | 1 | 26.2887 |
A | 103 | ALA | 1 | 5.81737 |
A | 104 | THR | 1 | 10.2041 |
A | 105 | GLN | 1 | 52.2302 |
A | 106 | CYS | 1 | 12.9647 |
A | 116 | THR | 1 | 19.9032 |
A | 117 | SER | 1 | 12.8987 |
A | 119 | THR | 1 | 22.7757 |
A | 121 | ILE | 1 | 15.189 |
A | 128 | TYR | 1 | 66.5253 |
A | 134 | THR | 1 | 50.6727 |
A | 135 | CYS | 1 | 2.61009 |
A | 136 | SER | 1 | 39.7118 |
A | 137 | LEU | 1 | 53.1561 |
A | 138 | ASN | 1 | 42.8299 |
A | 139 | ARG | 1 | 81.9165 |
A | 140 | TYR | 1 | 86.4117 |
A | 145 | TYR | 1 | 9.02516 |
A | 146 | GLY | 1 | 8.25783 |
A | 147 | PRO | 1 | 8.5554 |
A | 190 | SER | 1 | 9.85719 |
A | 191 | GLY | 1 | 15.6605 |
A | 192 | GLU | 1 | 2.93553 |
A | 241 | THR | 1 | 15.8578 |
A | 242 | ARG | 1 | 38.0833 |
A | 243 | VAL | 1 | 32.7166 |
A | 279 | TYR | 1 | 46.6128 |
A | 281 | SER | 1 | 43.2088 |
A | 282 | VAL | 1 | 11.9861 |
A | 283 | THR | 1 | 30.5658 |
A | 295 | ARG | 1 | 54.2229 |
B | 1 | GLN | 1 | 97.4462 |
B | 2 | VAL | 1 | 13.6402 |
B | 3 | GLN | 1 | 56.1574 |
B | 5 | GLN | 1 | 50.7106 |
B | 6 | GLU | 1 | 0.853008 |
B | 7 | SER | 1 | 8.57829 |
B | 26 | GLY | 1 | 24.476 |
B | 27 | SER | 1 | 23.1216 |
B | 28 | ILE | 1 | 2.73042 |
B | 29 | TYR | 1 | 99.7546 |
B | 30 | SER | 1 | 13.7472 |
B | 45 | HIS | 1 | 17.0975 |
B | 99 | SER | 1 | 14.3742 |
B | 100 | ASP | 1 | 49.1928 |
B | 101 | ARG | 1 | 109.806 |
B | 102 | LEU | 1 | 13.9772 |
B | 103 | THR | 1 | 20.9277 |
B | 104 | ASP | 1 | 6.27695 |
B | 108 | CYS | 1 | 16.5484 |
B | 109 | GLU | 1 | 50.0274 |
B | 110 | ALA | 1 | 7.51728 |
B | 111 | ASP | 1 | 47.5298 |
B | 112 | TYR | 1 | 43.1371 |
B | 113 | TRP | 1 | 71.6285 |
B | 114 | GLY | 1 | 3.13616 |
B | 115 | GLN | 1 | 67.258 |
To prepare data for assessment is to join a prediction output table with its corresponding weighted ground truth table.
For missing rows in the prediction output table, set the "prediction" value to 0.
Example of a combined predictions and weighted ground truth table:
chain | seqNum | resName | prediction | fact | weight |
A | 1 | ALA | 0 | 0 | 0 |
A | 2 | ILE | 0 | 0 | 0 |
A | 3 | ASN | 0 | 0 | 0 |
A | 4 | ALA | 0.95 | 1 | 15.4 |
A | 5 | ALA | 0.91 | 1 | 11.2 |
A | 6 | ARG | 0 | 1 | 0.7 |
A | 7 | GLU | 0 | 0 | 0 |
A | 8 | ASN | 0 | 0 | 0 |
A | 9 | LEU | 0.59 | 0 | 0 |
A | 10 | GLY | 0.15 | 1 | 1.2 |
A | 11 | ALA | 0.17 | 0 | 0 |
A | 12 | SER | 0 | 0 | 0 |
B | 1 | MET | 0 | 0 | 0 |
B | 2 | ALA | 0 | 0 | 0 |
B | 3 | GLN | 0.87 | 1 | 30.2 |
B | 4 | VAL | 0.21 | 1 | 34.5 |
B | 5 | GLN | 0.74 | 1 | 18.4 |
B | 6 | LEU | 0 | 1 | 8.2 |
B | 7 | GLN | 0 | 0 | 0 |
B | 8 | GLU | 0.47 | 1 | 0.4 |
B | 9 | SER | 0.14 | 0 | 0 |
B | 10 | GLY | 0 | 0 | 0 |
B | 11 | GLY | 0 | 0 | 0 |
B | 12 | GLY | 0 | 0 | 0 |
The SEPIQ Track A challenge is a binary classification challenge. The methodologies for evaluating binary classification results are well known.
As per-residue predictions will be provided as real values from 0 to 1, a threshold can be applied to convert these real-valued predictions into integer predictions (either 0 or 1). The integer predictions can then be compared to the "fact" column values of the ground truth.
All comparisons can be summarized using four values:
These values can be used to calculate many different assessment measures, such as accuracy, balanced accuracy, F1 score, precision, recall, specificity, MCC, and the Jaccard index. Values from a confusion matrix can be calculated for every valid threshold and used to plot a curve. The area under the resulting curve summarizes the overall performance of a classifier.
The most popular types of such curves are:
In antibody epitope prediction, target proteins typically consist of hundreds to thousands of amino acid residues, while the epitope comprises only about 10–20 residues. This creates a severe class imbalance, therefore we should consider several issues:
The weights are non-zero only in cases where fact = 1 (i.e., the residue participates in the interface). Therefore, the weights can be naturally applied to TP and FN values.
This means that we can define a weighted version of the recall measure.
The standard recall formula is:
where TP is the number of correctly recovered interface residues and (TP + FN) is the total number of interface residues.
However, we can replace:
Then the area-weighted recall formula becomes:
When we use the weighted recall together with the ordinary precision, we can calculate a weighted (area-informed) variation of the precision–recall curve and the corresponding area under the curve. Since PR-AUC (precision–recall area under the curve) is known to perform well with imbalanced data, two versions of PR-AUC (with and without weights) will be the primary evaluation measures for SEPIQ Track A.
The macro-average of the individual PR-AUCs of each antigen-VHH complex will be calculated. Since antibodies with larger epitope counts dominate, residues will not be pooled across all antibodies for a single PR-AUC calculation (micro-averaging).
Because the held-out structural test set is limited in size, final rankings will be accompanied by per-complex metrics and uncertainty estimates. If the evaluation metrics for the participants are closely matched, a resampling-based cluster bootstrap test will be performed to assess ranking robustness and reduce the impact of outliers.
Amino acids:
Ground Truth:
Team A Predictions:
Team B Predictions:
Step 1: Compute Observed Scores
Step 2: Bootstrap Resampling (B iterations, e.g., B = 2,000)
For b = 1, 2, . . . ,B:
Draw L indices randomly with replacement from {1, 2, . . . , L} to create index array Ib.
Use the exact same index set Ib across all antibodies and both teams to preserve pairing.
Form bootstrapped subset: .
Recompute PR-AUC for each antibody k:
Recompute Macro-Averages: .
Step 3: Statistical Significance Evaluation
Using the distribution of differences {}:
Sort Score(b) and locate the 2.5th and 97.5th percentiles.
Decision Rule: If the 95% Confidence Interval does NOT contain 0 (e.g., [0.012, 0.045]), the score difference between Team A and Team B is statistically significant (p<0.05).
The SEPIQ Track B challenge is a regression task with the goal of predicting structures of antigen-antibody complexes at the atomic level. Track B submissions will be evaluated through the established CAPRI assessment framework, including CAPRI-Q/DockQ-based measures. After optimal superimposition of predicted and ground-truth structures, the scoring procedure computes three key metrics: the root mean square deviation of ligand backbone atoms (L-RMSD), the backbone RMSD of all interface residues (i-RMSDbb), and the fraction of correctly predicted ligand–receptor contacts (fnat). These metrics are combined into a single composite score, DockQ, which ranges from 0 to 1 (see equation 1). Models with DockQ ≥ 0.80 are classified as high accuracy, while those with DockQ ≤ 0.23 are considered incorrect [Basu and Wallner, 2016]. The CAPRI-Q scoring protocol details are published by Collins et al. [2024]. The DockQ score and the CAPRI-Q scoring protocol offer a reliable and effective framework for evaluating predicted antigen-antibody atomic structures.
Equation 1: Calculation of the continuous quality score which combines L-RMSD, i-RMSDbb, and fnat metrics. The scaling parameters d1 and d2 determine how fast large RMSD values can be scaled to zero ().