Guideline

Abstract

Antigen-antibody interaction prediction remains a fundamental challenge in computational biology and AI-driven drug discovery. While recent machine learning approaches have enabled accurate single-chain protein structure prediction, modeling antigen-antibody interactions remains challenging because of extreme sequence diversity, evolutionary asymmetry, conformational flexibility, and, critically, the lack of large experimental datasets with functional labels and high-resolution structural supervision tailored to antigen-antibody interactions. As a result, existing models often fail to generalize beyond narrow benchmarks and curated structural complexes. We announce SEPIQ (Sequence-based Epitope Prediction Intelligent Query), a community challenge built around a large experimental dataset of immune-derived VHH nanobody sequences. SEPIQ comprises two complementary tracks: Track A, managed by SEPIQ, where participants submit per-antigen-residue epitope/paratope probabilities, and Track B, managed by CAPRI (Critical Assessment of PRediction of Interactions), where participants submit predicted antigen-VHH complex structures according to CAPRI procedures. A key strength of SEPIQ is the availability of high-resolution Cryo-EM ground truth for blind evaluation with a large-scale antigen-VHH interaction dataset. To ensure rigor, Cryo-EM structures and derived test labels will be held out during the challenge and released only after completion of the submission evaluation. The SEPIQ challenge aims to catalyze “Sequence-to-Structure Immunology” by testing a “folding from binding” hypothesis: large-scale binding-derived sequence data may contain statistical signatures of structural and functional constraints when combined with limited high-resolution structural ground truth.

Introduction

Accurate epitope identification and prediction of antigen-antibody interactions are critical for drug discovery, as they enable therapeutic design, define mechanisms of action, and allow assessment of efficacy and potential off-target effects. Despite their importance, reliable prediction of antigen-antibody interactions remains a central and unresolved challenge in computational biology. Recent advancements in machine learning led to the development of methods that excel at single-chain protein structure prediction, yet they remain limited for antigen-antibody modeling due to repertoire-scale sequence diversity, evolutionary asymmetry at interfaces, limited cross-chain coevolutionary signal in MSAs, epitope conformational variability, and binder/non-binder discrimination across large experimental panels. As a result, state-of-the-art models (e.g., AlphaFold3, OpenFold3, Boltz-1, and Chai-1) achieve acceptable accuracy on only a minority of cases in antigen-antibody docking benchmarks (40%), in contrast to their much higher success rates on single-chain protein structure prediction (≈90%) [The OpenFold3 Team, 2025, Yin and Pierce, 2024, Gaudreault et al., 2025, Hitawala and Gray, 2025, Abramson et al., 2024, Discovery et al., 2024, Jumper et al., 2021]. One of the key limiting factors in improving model performance for antigen-antibody complexes is the scarcity of experimental data that provides sequence diversity, functional binding outcomes, and high-resolution structural ground truth. The two largest publicly available datasets of antigen-antibody complex structures, SAbDab and SAAINT-DB, are curated from entries in the Protein Data Bank (PDB) [Huang et al., 2025, Schneider et al., 2022, Dunbar et al., 2014]. According to SAbDab, as of July 2026, the PDB contained 9,775 antibody structures with antigens, which underscores the limited scale of structural supervision available for predicting antigen-antibody complexes. Antigen-VHH complexes are even rarer (SAbDab-nano reports 2,013 structures with antigen). To accelerate the progress in prediction of antigen-antibody interactions, we announce the open SEPIQ machine learning challenge and provide novel large-scale experimental sequence data to enable robust modeling of antigen-antibody complexes, together with high-quality structural ground truth for evaluation, with test labels held out for blind assessment.

The competition will feature two primary tracks at different resolutions: Track A (residue-level resolution; SEPIQ) asks participants to develop sequence-based models that predict epitope/paratope residues (as per-residue probabilities or binary labels) from VHH and antigen sequence inputs, enabling interface characterization even without explicit 3D complex modeling. Track B (atomic-level resolution; CAPRI) asks participants to predict antigen-antibody complex structures from VHH and antigen sequences. We believe that the SEPIQ challenge will encourage the development of novel sequence-based architectures and has the potential to establish a new research subfield at the intersection of machine learning, immunology, and structural biology, accelerating the discovery of new medicines. Based on the social media outreach and positive feedback from in-person presentations, we expect more than 100 participants from diverse backgrounds to take part in the SEPIQ challenge per a year. Conceptually, SEPIQ tests a “folding from binding” hypothesis: large-scale sequence-level binding data may provide weak but scalable information about structural and functional constraints at antigen-antibody interfaces. We present this as a testable hypothesis rather than a definitive claim, because high-resolution structural labels remain limited and are used only for blind evaluation.

Tasks and application scenarios

This challenge involves antigen-antibody complexes containing novel epitopes and paratopes. Accurate prediction of epitope-paratope interactions can accelerate therapeutic nanobody discovery by reducing the need for costly wet-lab screening and structural characterization. Sequence-only methods are particularly valuable in settings where experimentally determined structures are unavailable.

Track A (residue-level resolution) asks participants to predict which residues on the antigen (epitope) and the antibody (paratope) form the binding interface of each antigen-VHH complex, using sequence input only. Compared to three-dimensional structural data, sequence-based data on antigen-antibody interactions can be acquired more extensively and at a lower cost. Such data will continue to accumulate at an accelerating rate. The primary objective of Track A is to explore the potential applications of these substantial sequence-based data on antigen-antibody interactions. However, there are no restrictions on the algorithms used to predict epitopes/paratopes for Track A. It is perfectly acceptable to predict epitopes/paratopes based on three-dimensional structure prediction (equivalent to Track B). The primary submission format is a per-residue epitope/paratope probability vector. Please refer to the section below for a more detailed explanation. AI models developed for Track A would be able to identify antigen-antibody interfaces directly from sequence data, enabling antibody design without experimental structural information. We demonstrate the possibility of solving the challenge with our lightweight baseline model, which is trained on a publicly available dataset, in the 'Baseline and code' section.

Track B (atomic-level resolution) asks participants to predict antigen-VHH complex structures from antigen and VHH sequences. This track will follow the established CAPRI framework for target release, model submission, and assessment. In this track, participants should submit predicted complex atomic structures. Please wait for a future announcement from CAPRI for a more detailed explanation. Although several existing models, including Boltz-1, Chai-1, and AlphaFold3, can predict antigen-antibody complex structures, their accuracy on this task remains limited. Improving these models would advance computational modeling of antigen-antibody interactions and accelerate drug development.

Although the tasks are highly challenging due to sparse experimental supervision and the complexity of immune recognition, recent advances in protein foundation models, geometric deep learning, and antibody structure prediction demonstrate that meaningful progress is achievable. The aim of this challenge is to make antibody drug discovery more efficient in the real world. As long as the answer is correct, we do not consider data leakage resulting from the use of publicly available data to be an issue.

Submission format for prediction output

For each complex sequence, provide a table with chain names (chain A for the antigen, chain B for the VHH), residue sequence numbers, and amino acid residue names.

Participants must return tables formed by adding a column to each input table in csv format. The added column must contain values from 0 to 1 indicating the probability that the residue is part of the binding interface.

Example of output for the antigen:

chain
seqNum
resName
prediction
A
1
ALA
0
A
2
ILE
0
A
3
ASN
0
A
4
ALA
0.95
A
5
ALA
0.91
A
6
ARG
0
A
7
GLU
0
A
8
ASN
0
A
9
LEU
0.59
A
10
GLY
0.15
A
11
ALA
0.17
A
12
SER
0
B
1
MET
0
B
2
ALA
0
B
3
GLN
0.87
B
4
VAL
0.21
B
5
GLN
0.74
B
6
LEU
0
B
7
GLN
0
B
8
GLU
0.47
B
9
SER
0.14
B
10
GLY
0
B
11
GLY
0
B
12
GLY
0

Participants may skip residues in their output — it will be assumed that the prediction is 0 for any omitted residue.

Example of shorter output for the antigen:

chain
seqNum
resName
prediction
A
4
ALA
0.95
A
5
ALA
0.91
A
9
LEU
0.59
A
10
GLY
0.15
A
11
ALA
0.17
B
3
GLN
0.87
B
4
VAL
0.21
B
5
GLN
0.74
B
8
GLU
0.47
B
9
SER
0.14

The template CSV files for predictions have been uploaded to the Hugging Face repository.

Ground truth generation

Given an experimentally determined antigen-VHH complex structure, the interface will be described using the Voronota-LT method, which constructs Voronoi tessellation-based inter-atom contact surfaces and calculates their areas.

We will adopt the Voronota-LT software [Kliment et al., 2025] to calculate only the relevant contacts — that is, the contacts between the antigen and the VHH chains. Moreover, we can instruct Voronota-LT to summarize (aggregate) contact areas for each residue that has at least one relevant contact.

Let us illustrate the interface determination process using a known antigen-VHH structure from the Protein Data Bank (PDB).

Firstly, let us obtain the software:

# download the universal Voronota-LT version 1.1.479 executable
wget "https://github.com/kliment-olechnovic/voronota/releases/download/v1.29.4602/cosmopolitan_voronota-lt_v1.1.479.exe"
# rename the downloaded executable
mv cosmopolitan_voronota-lt_v1.1.479.exe voronota-lt
# set execute permission
chmod +x ./voronota-lt

Secondly, let us prepare an input structure — download and unzip the first assembly structure from the PDB entry 5F7L (https://www.rcsb.org/structure/5F7L):

wget "https://files.rcsb.org/download/5F7L-assembly1.cif.gz"
gunzip ./5F7L-assembly1.cif.gz

We note that in the downloaded structure there are two chains: chain A is the antigen, and chain B is the VHH.

Finally, let us run Voronota-LT to produce a table of residue-level binding site areas:

./voronota-lt --input "./5F7L-assembly1.cif" \
  --restrict-contacts '[-a1 [-chain A] -a2 [-chain B]]' \
  --write-sites-residue-level-to-file ./sites-residue-level.tsv

After that, the generated file sites-residue-level.tsv contains the binding site areas for all residues that participate in the interface:

rs_header
ID_chain
ID_rnum
ID_rname
area
arc_length
distance
count
rs
A
102
TYR
26.2887
5.35143
3.25509
10
rs
A
103
ALA
5.81737
2.34553
3.9238
3
rs
A
104
THR
10.2041
1.92648
3.50697
3
rs
A
105
GLN
52.2302
6.67633
3.06804
31
rs
A
106
CYS
12.9647
0
2.69348
6
rs
A
116
THR
19.9032
6.73551
4.12098
13
rs
A
117
SER
12.8987
15.7238
4.87017
9
rs
A
119
THR
22.7757
11.912
3.51999
10
rs
A
121
ILE
15.189
7.36034
4.00473
3
rs
A
128
TYR
66.5253
20.6576
3.23458
34
rs
A
134
THR
50.6727
14.9325
2.83448
24
rs
A
135
CYS
2.61009
0
3.58477
3
rs
A
136
SER
39.7118
3.33127
2.52272
25
rs
A
137
LEU
53.1561
0
3.44333
35
rs
A
138
ASN
42.8299
5.52968
2.7755
25
rs
A
139
ARG
81.9165
19.6108
3.05166
43
rs
A
140
TYR
86.4117
17.5458
2.50269
39
rs
A
145
TYR
9.02516
5.71504
4.12956
5
rs
A
146
GLY
8.25783
2.21431
4.67829
4
rs
A
147
PRO
8.5554
0.783521
4.72983
5
rs
A
190
SER
9.85719
4.90913
4.31282
4
rs
A
191
GLY
15.6605
1.62435
4.07761
5
rs
A
192
GLU
2.93553
4.35852
5.05939
2
rs
A
241
THR
15.8578
6.29968
4.28125
7
rs
A
242
ARG
38.0833
24.4135
4.69196
14
rs
A
243
VAL
32.7166
18.5909
3.30051
9
rs
A
279
TYR
46.6128
3.07466
2.54624
28
rs
A
281
SER
43.2088
8.06525
2.49241
15
rs
A
282
VAL
11.9861
3.03489
3.59902
9
rs
A
283
THR
30.5658
7.36819
3.50735
16
rs
A
295
ARG
54.2229
21.4632
3.40346
24
rs
B
1
GLN
97.4462
18.6789
2.66855
57
rs
B
2
VAL
13.6402
2.62611
2.52272
8
rs
B
3
GLN
56.1574
21.8633
3.51999
19
rs
B
5
GLN
50.7106
26.0035
4.28125
22
rs
B
6
GLU
0.853008
1.11928
5.77592
1
rs
B
7
SER
8.57829
8.99021
5.43684
1
rs
B
26
GLY
24.476
11.6387
3.23458
7
rs
B
27
SER
23.1216
3.52698
3.94957
11
rs
B
28
ILE
2.73042
1.68464
4.77136
2
rs
B
29
TYR
99.7546
17.6416
2.49241
54
rs
B
30
SER
13.7472
6.48282
3.24403
5
rs
B
45
HIS
17.0975
8.7078
4.61964
9
rs
B
99
SER
14.3742
0
3.19291
4
rs
B
100
ASP
49.1928
8.82601
2.54624
23
rs
B
101
ARG
109.806
20.2009
2.97931
65
rs
B
102
LEU
13.9772
0.810522
3.74136
7
rs
B
103
THR
20.9277
7.30925
3.96832
9
rs
B
104
ASP
6.27695
6.1968
4.84443
2
rs
B
108
CYS
16.5484
6.77437
3.23494
5
rs
B
109
GLU
50.0274
12.2063
2.50269
19
rs
B
110
ALA
7.51728
0
3.85594
5
rs
B
111
ASP
47.5298
0
2.7755
28
rs
B
112
TYR
43.1371
1.95946
3.34279
27
rs
B
113
TRP
71.6285
19.3995
3.60848
46
rs
B
114
GLY
3.13616
5.68057
5.47438
3
rs
B
115
GLN
67.2589
33.2266
3.30051
24

From that table, we extract only the required columns (ID_chain, ID_rnum, ID_rname, area), rename them to (chain, seqNum, resName, weight) respectively, insert an additional fact column with all values set to 1 (because all reported residues are interface residues), and generate a table corresponding to the ground truth format:

{
echo "chain seqNum resName fact weight" | sed 's/\s\+/\t/g'
cat "./sites-residue-level.tsv" | tail -n +2 | awk '{print $2 "\t" $3 "\t" $4 "\t1\t" $5}'
} \
> "./positive_ground_truth.tsv"

This will produce a table of positive ground truth:

chain
seqNum
resName
fact
weight
A
102
TYR
1
26.2887
A
103
ALA
1
5.81737
A
104
THR
1
10.2041
A
105
GLN
1
52.2302
A
106
CYS
1
12.9647
A
116
THR
1
19.9032
A
117
SER
1
12.8987
A
119
THR
1
22.7757
A
121
ILE
1
15.189
A
128
TYR
1
66.5253
A
134
THR
1
50.6727
A
135
CYS
1
2.61009
A
136
SER
1
39.7118
A
137
LEU
1
53.1561
A
138
ASN
1
42.8299
A
139
ARG
1
81.9165
A
140
TYR
1
86.4117
A
145
TYR
1
9.02516
A
146
GLY
1
8.25783
A
147
PRO
1
8.5554
A
190
SER
1
9.85719
A
191
GLY
1
15.6605
A
192
GLU
1
2.93553
A
241
THR
1
15.8578
A
242
ARG
1
38.0833
A
243
VAL
1
32.7166
A
279
TYR
1
46.6128
A
281
SER
1
43.2088
A
282
VAL
1
11.9861
A
283
THR
1
30.5658
A
295
ARG
1
54.2229
B
1
GLN
1
97.4462
B
2
VAL
1
13.6402
B
3
GLN
1
56.1574
B
5
GLN
1
50.7106
B
6
GLU
1
0.853008
B
7
SER
1
8.57829
B
26
GLY
1
24.476
B
27
SER
1
23.1216
B
28
ILE
1
2.73042
B
29
TYR
1
99.7546
B
30
SER
1
13.7472
B
45
HIS
1
17.0975
B
99
SER
1
14.3742
B
100
ASP
1
49.1928
B
101
ARG
1
109.806
B
102
LEU
1
13.9772
B
103
THR
1
20.9277
B
104
ASP
1
6.27695
B
108
CYS
1
16.5484
B
109
GLU
1
50.0274
B
110
ALA
1
7.51728
B
111
ASP
1
47.5298
B
112
TYR
1
43.1371
B
113
TRP
1
71.6285
B
114
GLY
1
3.13616
B
115
GLN
1
67.258

An assessment-ready format of combined predictions and ground truth data

To prepare data for assessment is to join a prediction output table with its corresponding weighted ground truth table.

For missing rows in the prediction output table, set the "prediction" value to 0.

Example of a combined predictions and weighted ground truth table:

chain
seqNum
resName
prediction
fact
weight
A
1
ALA
0
0
0
A
2
ILE
0
0
0
A
3
ASN
0
0
0
A
4
ALA
0.95
1
15.4
A
5
ALA
0.91
1
11.2
A
6
ARG
0
1
0.7
A
7
GLU
0
0
0
A
8
ASN
0
0
0
A
9
LEU
0.59
0
0
A
10
GLY
0.15
1
1.2
A
11
ALA
0.17
0
0
A
12
SER
0
0
0
B
1
MET
0
0
0
B
2
ALA
0
0
0
B
3
GLN
0.87
1
30.2
B
4
VAL
0.21
1
34.5
B
5
GLN
0.74
1
18.4
B
6
LEU
0
1
8.2
B
7
GLN
0
0
0
B
8
GLU
0.47
1
0.4
B
9
SER
0.14
0
0
B
10
GLY
0
0
0
B
11
GLY
0
0
0
B
12
GLY
0
0
0

Predictor’s performance evaluation

General considerations

The SEPIQ Track A challenge is a binary classification challenge. The methodologies for evaluating binary classification results are well known.

As per-residue predictions will be provided as real values from 0 to 1, a threshold can be applied to convert these real-valued predictions into integer predictions (either 0 or 1). The integer predictions can then be compared to the "fact" column values of the ground truth.

All comparisons can be summarized using four values:

  • TP = the number of true positive predictions, i.e., the number of cases where fact = 1 and the integer prediction = 1
  • FN = the number of false negative predictions, i.e., the number of cases where fact = 1 and the integer prediction = 0
  • TN = the number of true negative predictions, i.e., the number of cases where fact = 0 and the integer prediction = 0
  • FP = the number of false positive predictions, i.e., the number of cases where fact = 0 and the integer prediction = 1

These values can be used to calculate many different assessment measures, such as accuracy, balanced accuracy, F1 score, precision, recall, specificity, MCC, and the Jaccard index. Values from a confusion matrix can be calculated for every valid threshold and used to plot a curve. The area under the resulting curve summarizes the overall performance of a classifier.

The most popular types of such curves are:

  • ROC (receiver operating characteristic) curve — represents the trade-off between the true positive rate (sensitivity) and the false positive rate (1 − specificity) across different thresholds.
  • PR (precision–recall) curve — illustrates the trade-off between precision and recall across different thresholds.

Measures of choice for SEPIQ

In antibody epitope prediction, target proteins typically consist of hundreds to thousands of amino acid residues, while the epitope comprises only about 10–20 residues. This creates a severe class imbalance, therefore we should consider several issues:

  • we may not want to fix the decision threshold for converting real-valued predictions to integer (0 or 1) predictions;
  • the ground truth data will likely be highly imbalanced — there will be many more non-interface residues than interface residues;
  • we want to use weights (per-residue binding site area) in the assessment to put more emphasis on residues that are more involved in the interface.

The weights are non-zero only in cases where fact = 1 (i.e., the residue participates in the interface). Therefore, the weights can be naturally applied to TP and FN values.

This means that we can define a weighted version of the recall measure.

The standard recall formula is:

recall=TPTP+FNrecall=\frac{TP}{TP+FN}

where TP is the number of correctly recovered interface residues and (TP + FN) is the total number of interface residues.

However, we can replace:

  • TP with TP_area — the total sum of the weights of the correctly recovered interface residues (i.e., the total recovered interface area), and
  • (TP + FN) with P_area — the total sum of the weights of all interface residues (i.e., the total ground truth interface area).

Then the area-weighted recall formula becomes:

recallweighted=TPareaParearecall_{weighted}=\frac{TP_{area}}{P_{area}}

When we use the weighted recall together with the ordinary precision, we can calculate a weighted (area-informed) variation of the precision–recall curve and the corresponding area under the curve. Since PR-AUC (precision–recall area under the curve) is known to perform well with imbalanced data, two versions of PR-AUC (with and without weights) will be the primary evaluation measures for SEPIQ Track A.

Ranking participants based on PR-AUC

The macro-average of the individual PR-AUCs of each antigen-VHH complex will be calculated. Since antibodies with larger epitope counts dominate, residues will not be pooled across all antibodies for a single PR-AUC calculation (micro-averaging).

Score=13∑k=13PR AUCk for Antibody k ∈ 1,2,3Score = \frac{1}{3}\sum_{k=1}^{3}PR\:AUC_k\:for\: Antibody\:k\:∈\:{1,2,3}

Because the held-out structural test set is limited in size, final rankings will be accompanied by per-complex metrics and uncertainty estimates. If the evaluation metrics for the participants are closely matched, a resampling-based cluster bootstrap test will be performed to assess ranking robustness and reduce the impact of outliers.

Amino acids: i∈1,2,3,...,L for Sequence Length Li ∈ {1, 2, 3, . . . , L}\: for\: Sequence\: Length\: L

Ground Truth: yk,i ∈ 0,1 for Antibody k ∈ 1,2,3y_{k,i}\: ∈\: {0,1}\: for\: Antibody\: k\: ∈\: {1,2,3}

Team A Predictions: pA,k,i∈[0,1]p_{A, k, i} ∈ [0, 1]

Team B Predictions: pB,k,i∈[0,1]p_{B, k, i} ∈ [0, 1]

Step 1: Compute Observed Scores

  • Calculate PR-AUC for each antibody k for both Team A and Team B:
PR AUCA,k=PR AUC(yk,pA,k)PR AUCB,k=PR AUC(yk,pB,k)\begin{gathered}PR\:AUC_{A, k}= PR\:AUC(y_k, p_{A, k})\\ PR\: AUC_{B, k}= PR\: AUC(y_k, p_{B, k})\end{gathered}
  • Compute Macro-Averages:
ScoreA=13∑k=13PR AUCA,k,  ScoreB=13∑k=13PR AUCB,kScore_A = \frac13\sum_{k=1}^{3}PR\: AUC_{A, k},\:\: Score_B = \frac13\sum_{k=1}^{3}PR\: AUC_{B, k}
  • Calculate observed delta:
ΔScoreobs=ScoreA−ScoreB\Delta Score_{obs} = Score_A-Score_B

Step 2: Bootstrap Resampling (B iterations, e.g., B = 2,000)

For b = 1, 2, . . . ,B:

  • Resample Residue Indices with Replacement

    Draw L indices randomly with replacement from {1, 2, . . . , L} to create index array Ib.

    Use the exact same index set Ib across all antibodies and both teams to preserve pairing.

  • Recompute Metrics on Resampled Data

    Form bootstrapped subset: yk,Ib,pA,k,Ib,and pB,k,Iby_{k, I_b}, p_{A, k, I_b}, and\: p_{B, k, I_b}.

    Recompute PR-AUC for each antibody k:

    PR AUCA,k(b)=PR AUC(yk,Ib,pA,k,Ib)PR AUCB,k(b)=PR AUC(yk,Ib,pB,k,Ib)\begin{gathered}PR\: AUC_{A, k}^{(b)} = PR\: AUC(y_{k, I_b}, p_{A, k, I_b})\\ PR\: AUC_{B, k}^{(b)} = PR\: AUC(y_{k, I_b}, p_{B, k, I_b})\end{gathered}

    Recompute Macro-Averages: ScoreA(b) and ScoreB(b)Score_A^{(b)} \:and\: Score_B^{(b)}.

  • Record Bootstrapped Difference:
    ΔScore(b)=ScoreA(b)−ScoreB(b)\Delta Score^{(b)} = Score_A^{(b)} - Score_B^{(b)}

Step 3: Statistical Significance Evaluation

Using the distribution of differences {ΔScore(1),...,ΔScore(B)\Delta Score^{(1)}, . . . , \Delta Score^{(B)}}:

  • 95% Confidence Interval (Percentile Method):

    Sort Score(b) and locate the 2.5th and 97.5th percentiles.

    Decision Rule: If the 95% Confidence Interval does NOT contain 0 (e.g., [0.012, 0.045]), the score difference between Team A and Team B is statistically significant (p<0.05).

  • Empirical p-value:
    p=2×min(∑b=1BI(ΔScore(b)≤0)B,∑b=1BI(ΔScore(b)≥0)B)p=2×min\Bigg(\frac{\sum_{b=1}^{B}I(\Delta Score^{(b)} ≤ 0)}{B}, \frac{ \sum_{b=1}^{B}I(\Delta Score^{(b)} ≥ 0)}{B}\Bigg)

Supplementary Information (For official information, please visit the CAPRI website)

The SEPIQ Track B challenge is a regression task with the goal of predicting structures of antigen-antibody complexes at the atomic level. Track B submissions will be evaluated through the established CAPRI assessment framework, including CAPRI-Q/DockQ-based measures. After optimal superimposition of predicted and ground-truth structures, the scoring procedure computes three key metrics: the root mean square deviation of ligand backbone atoms (L-RMSD), the backbone RMSD of all interface residues (i-RMSDbb), and the fraction of correctly predicted ligand–receptor contacts (fnat). These metrics are combined into a single composite score, DockQ, which ranges from 0 to 1 (see equation 1). Models with DockQ ≥ 0.80 are classified as high accuracy, while those with DockQ ≤ 0.23 are considered incorrect [Basu and Wallner, 2016]. The CAPRI-Q scoring protocol details are published by Collins et al. [2024]. The DockQ score and the CAPRI-Q scoring protocol offer a reliable and effective framework for evaluating predicted antigen-antibody atomic structures.

DockQ=13(fnat+11+(L−RMSDd1)2+11+(i−RMSDbbd2)2)DockQ = \frac{1}{3}\Bigg(f_{nat} + \frac{1}{1 + (\frac{L-RMSD}{d_1})^2} + \frac{1}{1+(\frac{i-RMSDbb}{d_2})^2}\Bigg)

Equation 1: Calculation of the continuous quality score which combines L-RMSD, i-RMSDbb, and fnat metrics. The scaling parameters d1 and d2 determine how fast large RMSD values can be scaled to zero (d1=8.5A˚ and d2=1.5A˚d_1 = 8.5 Å \:and\: d_2 = 1.5 Å).