SEPIQ 2026 Challenge targets, timeline, and dataset

Targets

HER2 is a clinically important oncology target for diagnosis and therapy. For this season, we have released a ∼ 1.8M-scale human HER2 VHH dataset, AVIDa-hHER2, with access to read-level NGS data (including singleton reads) with a subset of sequences labeled as binder/non-binder using statistical analysis. While the large-scale HER2 VHH sequence dataset provides useful information on binding behavior at the sequence level, it does not directly include residue-level epitope annotations. As a high-resolution ground truth for blind evaluation, three HER2-VHH complexes selected from the HER2 repertoire have been resolved (each defining a distinct HER2-VHH interface) to provide atomically detailed interface annotations. These structures and all derived labels will be held out during the challenge and released only after completion.

For submission, the human HER2 residues should be numbered according to the full-length UniProt P04626 sequence, including the signal peptide. The HER2 antigen used in this experiment is an extracellular domain which may form dimers under experimental conditions. The full amino acid sequence of the HER2 antigen used in this experiment is described below:

MELAALCRWG10  LLLALLPPGA20  ASTQVCTGTD30  MKLRLPASPE40  THLDMLRHLY50  QGCQVVQGNL60  ELTYLPTNAS70  LSFLQDIQEV80  QGYVLIAHNQ90  VRQVPLQRLR100  IVRGTQLFED110  NYALAVLDNG120  DPLNNTTPVT130  GASPGGLREL140  QLRSLTEILK150  GGVLIQRNPQ160  LCYQDTILWK170  DIFHKNNQLA180  LTLIDTNRSR190  ACHPCSPMCK200  GSRCWGESSE210  DCQSLTRTVC220  AGGCARCKGP230  LPTDCCHEQC240  AAGCTGPKHS250  DCLACLHFNH260  SGICELHCPA270  LVTYNTDTFE280  SMPNPEGRYT290  FGASCVTACP300  YNYLSTDVGS310  CTLVCPLHNQ320  EVTAEDGTQR330  CEKCSKPCAR340  VCYGLGMEHL350  REVRAVTSAN360  IQEFAGCKKI370  FGSLAFLPES380  FDGDPASNTA390  PLQPEQLQVF400  ETLEEITGYL410  YISAWPDSLP420  DLSVFQNLQV430  IRGRILHNGA440  YSLTLQGLGI450  SWLGLRSLRE460  LGSGLALIHH470  NTHLCFVHTV480  PWDQLFRNPH490  QALLHTANRP500  EDECVGEGLA510  CHQLCARGHC520  WGPGPTQCVN530  CSQFLRGQEC540  VEECRVLQGL550  PREYVNARHC560  LPCHPECQPQ570  NGSVTCFGPE580  ADQCVACAHY590  KDPPFCVARC600  PSGVKPDLSY610  MPIWKFPDEE620      GACQ      HHHHHH\begin{gathered}MELAALCRWG^{10}\:\: LLLALLPPGA^{20}\:\: ASTQVCTGTD^{30}\:\: MKLRLPASPE^{40}\:\: \\ THLDMLRHLY^{50}\:\: QGCQVVQGNL^{60}\:\: ELTYLPTNAS^{70}\:\: LSFLQDIQEV^{80}\:\: \\ QGYVLIAHNQ^{90}\:\: VRQVPLQRLR^{100}\:\: IVRGTQLFED^{110}\:\: NYALAVLDNG^{120}\:\: \\ DPLNNTTPVT^{130}\:\: GASPGGLREL^{140}\:\: QLRSLTEILK^{150}\:\: GGVLIQRNPQ^{160}\:\: \\ LCYQDTILWK^{170}\:\: DIFHKNNQLA^{180}\:\: LTLIDTNRSR^{190}\:\: ACHPCSPMCK^{200}\:\: \\ GSRCWGESSE^{210}\:\: DCQSLTRTVC^{220}\:\: AGGCARCKGP^{230}\:\: LPTDCCHEQC^{240}\:\: \\ AAGCTGPKHS^{250}\:\: DCLACLHFNH^{260}\:\: SGICELHCPA^{270}\:\: LVTYNTDTFE^{280}\:\: \\ SMPNPEGRYT^{290}\:\: FGASCVTACP^{300}\:\: YNYLSTDVGS^{310}\:\: CTLVCPLHNQ^{320}\:\: \\ EVTAEDGTQR^{330}\:\: CEKCSKPCAR^{340}\:\: VCYGLGMEHL^{350}\:\: REVRAVTSAN^{360}\:\: \\ IQEFAGCKKI^{370}\:\: FGSLAFLPES^{380}\:\: FDGDPASNTA^{390}\:\: PLQPEQLQVF^{400}\:\: \\ ETLEEITGYL^{410}\:\: YISAWPDSLP^{420}\:\: DLSVFQNLQV^{430}\:\: IRGRILHNGA^{440}\:\: \\ YSLTLQGLGI^{450}\:\: SWLGLRSLRE^{460}\:\: LGSGLALIHH^{470}\:\: NTHLCFVHTV^{480}\:\: \\ PWDQLFRNPH^{490}\:\: QALLHTANRP^{500}\:\: EDECVGEGLA^{510}\:\: CHQLCARGHC^{520}\:\: \\ WGPGPTQCVN^{530}\:\: CSQFLRGQEC^{540}\:\: VEECRVLQGL^{550}\:\: PREYVNARHC^{560}\:\: \\ LPCHPECQPQ^{570}\:\: NGSVTCFGPE^{580}\:\: ADQCVACAHY^{590}\:\: KDPPFCVARC^{600}\:\: \\ PSGVKPDLSY^{610}\:\: MPIWKFPDEE^{620}\:\:\:\:\:\:GACQ\:\:\:\:\:\:HHHHHH\end{gathered}

The residues of each VHH clone should be numbered according to sequential indexing. This season, we are not adopting the Kabat, Chothia, or IMGT schemes. The sequences of the three VHH clones for the challenge are listed below:

>ONU11

MAQVQLQESG10  GGSVQAGGSL20  RLSCAASGRT30  LSGWVMGWFR40  QAPGKDRDFV50  AHLSGDTTAY60  AHSVKGRFTI70  SRDNAKDTVF80  LQMNSLKPED90  TGVYYCAARP100  LGSTWTEHPY110  WGQGTQVTVS120      S    HHHHHH\begin{gathered}MAQVQLQESG^{10}\:\: GGSVQAGGSL^{20}\:\: RLSCAASGRT^{30}\:\: LSGWVMGWFR^{40}\:\: \\ QAPGKDRDFV^{50}\:\: AHLSGDTTAY^{60}\:\: AHSVKGRFTI^{70}\:\: SRDNAKDTVF^{80}\:\: \\ LQMNSLKPED^{90}\:\: TGVYYCAARP^{100}\:\: LGSTWTEHPY^{110}\:\: WGQGTQVTVS^{120}\:\: \\ \:\:\:\:S\:\:\:\: HHHHHH\end{gathered}

>ONU13

MAQVQLQESG10  GGLVQAGGSL20  RLSCTASERT30  FSIDYMGWYR40  QAPGKQREFV50  GTIEWNSGST60  GYGDSVKGRF70  TISRDNAKNT80  GYLQMNSLKP90  EDTAVYYCAA100  KYGSSRLWDY110  WGQGTQVTVS120  S    HHHHHH\begin{gathered}MAQVQLQESG^{10}\:\: GGLVQAGGSL^{20}\:\: RLSCTASERT^{30}\:\: FSIDYMGWYR^{40}\:\: \\ QAPGKQREFV^{50}\:\: GTIEWNSGST^{60}\:\: GYGDSVKGRF^{70}\:\: TISRDNAKNT^{80}\:\: \\ GYLQMNSLKP^{90}\:\: EDTAVYYCAA^{100}\:\: KYGSSRLWDY^{110}\:\: WGQGTQVTVS^{120}\:\: \\ S\:\:\:\:HHHHHH\end{gathered}

>ONU23

MAQVQLQESG10  GGLVQAGGSL20  RLSCAASGRT30  LSGAAMGWFR40  QAPGKEREFV50  AGISWSTSRT60  QYADSVKGRF70  TISRDNAKNT80  VFLQMNSLNA90  EDTAVYYCAA100  DGRFYSDYVW110  SSPDEYAYWG120  QGTQVTVSSH130  HHHHH\begin{gathered}MAQVQLQESG^{10}\:\: GGLVQAGGSL^{20}\:\: RLSCAASGRT^{30}\:\: LSGAAMGWFR^{40}\:\: \\ QAPGKEREFV^{50}\:\: AGISWSTSRT^{60}\:\: QYADSVKGRF^{70}\:\: TISRDNAKNT^{80}\:\: \\ VFLQMNSLNA^{90}\:\: EDTAVYYCAA^{100}\:\: DGRFYSDYVW^{110}\:\: SSPDEYAYWG^{120}\:\: \\ QGTQVTVSSH^{130}\:\: HHHHH\end{gathered}

Timeline

Challenge timeline: submission (Sep 1–Oct 31, 2026), evaluation and results announcement (Dec 15, 2026). Organizers will extend the duration of the challenge in case of any delays or technical issues from the SEPIQ or Hugging Face side.

Dataset generation

For the SEPIQ 2026 Competition, we constructed labeled immune-derived VHH repertoires following published protocols (IL-6 datasets in [Tsuruta et al., 2023]). Library Construction and Selection: One alpaca was immunized with wild-type (WT) human HER2. We generated two VHH libraries from blood samples via phage display [Bazan et al., 2012, Smith, 1985] using the pMES4 vector and performed one round of biopanning against HER2. Sequencing and Preprocessing: We performed next-generation sequencing (NGS; 105 paired reads per library), removed low-quality reads, translated DNA sequences to amino acids, and represented each unique VHH by its read count. Statistical Labeling: To identify binders, we compared VHH proportions before and after panning. We modeled the proportion change using the chi-squared test. For each VHH-target pair, we calculated p-values for each sublibrary in relation to the corresponding mother library. A VHH-target pair was labeled a "binder" if the VHH’s proportion increased upon panning relative to the corresponding mother library and the p-value was ≤ 0.05. Because these labels are derived from panning enrichment rather than direct biophysical measurements, we treat them as enrichment-derived binder/non-binder labels and will provide multiple-testing-aware metadata. Cryo-EM Structure Determination and Interface Annotation: Structural Cryo-EM data were obtained for three HER2-VHH complexes generated through the BINDS program. Protein complexes were vitrified, and three-dimensional reconstructions were obtained from aligned two-dimensional particle images. For each held-out complex, structure-quality indicators will be reported, including map resolution, model-validation statistics, and the confidence of interface annotations based on the fitted atomic model and local map density.

VHH dataset and Cryo-EM ground truth

The AVIDa-hHER2 dataset includes: 1) Experimental Labels: A few hundred sequences can be labeled as binder/non-binder using statistical analysis, providing a supervised signal where available. 2) Sequence Diversity: The data captures immune-derived diversity across multiple sampling time points, reflecting the natural maturation of antibody responses. 3) Read-level resource: NGS read-level data (including singleton reads) can support representation learning and robustness analyses. The data has been uploaded as 'NGS_read_count_table.csv'. For more detailed information, please refer to the description on the Hugging Face. The held-out Cryo-EM ground truth will include: 1) Atomic-Level Annotations: Epitope and paratope residue identities are derived from atomic models refined against Cryo-EM density maps and used to define residue-level interface annotations. 2) Contact Mapping: Interface contact maps derived from the refined atomic models, defining the physical constraints of the antigen-antibody interaction. 3) Structural Diversity: Multiple HER2-VHH complexes enabling analysis of distinct binding modes and epitope usage. SEPIQ has permission to use all of the above datasets collected and provided by SEPIQ coorganizers: COGNANO Inc., The University of Osaka, and the BINDS program. All animal experiments on the alpacas were conducted in accordance with the KYODOKEN Institute for Animal Science Research and Development (Kyoto, Japan) and the ARRIVE (Animal Research: Reporting of In Vivo Experiments) guidelines.

Baseline and code

Recently it has been shown that by training a transformer-based AntiBinder model on sequence data only with global binding labels and analyzing learned attention weights enables identification of epitope and paratope residues [Zhang et al., 2025]. As a baseline feasibility demonstration, we examined the AntiBinder model. Since the test set HER2 structures are held private until the end of the competition, to avoid data leakage, we train and evaluate the baseline models on similar publicly available data. As the ground truth antigen-antibody structure, we use the neutralizing VHH P17 in complex with SARS-CoV-2 spike receptor-binding domain (PDB ID: 8GZ5). For training, we use the public database of antigen-antibody interactions SAbDab-nano [Schneider et al., 2022], and the AVIDa-SARS-CoV-2 dataset [Tsuruta et al., 2023] released by COGNANO Inc. Both training datasets contain antibody and antigen sequence data, and binder/non-binder labels. The model trained on SAbDab-nano alone failed to predict binding for the 8GZ5 complex, resulting in epitope prediction performance comparable to the random baseline. In contrast, the model pretrained on AVIDa-SARS-CoV-2 and subsequently fine-tuned on SAbDab-nano correctly classified the complex as binding and achieved improved performance (Figure 2a). The best performance, corresponding to a PR-AUC of 0.176, is modest but above the expected performance of a random predictor. For 8GZ5, the fraction of epitope residues is 26/206 = 0.126, which provides an approximate expected PR-AUC for a random predictor. Considering the simplicity of the baseline method, these results suggest that pretraining on sequence-level binding data can provide useful signals for epitope prediction.

Image

Figure 1: Performance of baseline models. a) The PR AUC scores achieved by random predictor and AntiBinder models trained using SAbDab-nano and AVIDa-SARS-CoV-2 sequence datasets.

To promote reproducibility and offer participants a clear starting point, we have released a lightweight, sequence-based baseline model. The Starter Kit is available as Jupyter Notebooks containing: 1) instructions on accessing the dataset on Hugging Face; 2) an example of simple pre-training/inference using the baseline model; 3) illustrative scripts for evaluating per-residue epitope predictions; and 4) instructions for training and fine-tuning common open-source models suitable for Track A. Training the AntiBinder model required ∼30 minutes on a single consumer-grade GPU (NVIDIA GeForce RTX 3090 Ti), demonstrating that the baseline is both fast and accessible.

Contact

If participants have any questions or concerns, the SEPIQ support team is available via support email (info@cognano.co.jp). Due to the global distribution of the SEPIQ team, our organizers are available 24/7 for participant support.

Organizational aspects

Protocol

We suggest the following steps for participation in Track A of the SEPIQ 2026 competition: 1) create a Hugging Face account and access the SEPIQ 2026 competition page; 2) carefully review the instructions on dataset access, submission format, evaluation, and allowed resources; 3) develop a model using the released training data and permitted public resources; and 4) upload Track A submissions to the SEPIQ/Hugging Face platform. The primary Track A submission is a per-antigen-residue epitope/paratope probability vector in CSV format. Track B will follow the standard CAPRI registration, target release, model submission, and assessment procedures. Track B submissions will therefore not be handled as standard SEPIQ/Hugging Face submissions unless otherwise agreed with the CAPRI team.

Rules and Engagement

  • Team limits: Maximum team size of 10.
  • Submission Limits: To reduce leaderboard probing, each team will be allowed up to five Track A submissions. Track B submissions will follow the standard CAPRI submission rules.
  • Eligibility: 1) (a) To participate in Track A, participants must have a registered Hugging Face account. Track B eligibility and registration will follow the standard CAPRI procedures; (b) 18 years old or have reached the age of majority in their jurisdiction. Exceptions are made only if the Competition Sponsor approves and receives the necessary parental/guardian permissions; (c) not a resident of Crimea, so-called Donetsk People’s Republic (DNR) or Luhansk People’s Republic (LNR), the Russian Federation, the Republic of Belarus, Cuba, Iran, Syria, or North Korea; and (d) an individual or a representative of an entity subject to trade sanctions or export controls imposed by Japan, Ukraine, the European Union, the United Kingdom, or the United States. 2) The Competition is open to residents worldwide, with the exception of those residing in the Russian Federation, the Republic of Belarus, Cuba, Iran, Syria, North Korea, or the occupied territories of Ukraine (including Crimea, the so-called Donetsk People’s Republic (DNR), and the so-called Luhansk People’s Republic (LNR)). Furthermore, individuals or entities subject to export controls or sanctions imposed by Japan, Ukraine, the European Union, the United Kingdom, or the United States are ineligible to enter. It is your responsibility to ensure that participation in a skills-based competition is legal under your local jurisdiction. The Competition Host reserves the right to withhold or provide alternative prizes to comply with local regulations. If a winner resides in a region where legal or logistical restrictions prevent the awarding of a prize, they will be deemed ineligible to receive it. 3) If you enter as a representative of a company, educational institution, or other legal entity, or on behalf of an employer, these rules bind both you individually and the entity you represent. By entering, you warrant that your employer or the entity has full knowledge of your participation and has provided explicit consent, including for the potential receipt of a Prize. You further guarantee that your participation does not violate any internal policies, non-disclosure agreements, or procedures of your employer or entity. 4) The Competition Sponsor reserves the right to verify the eligibility of any participant and to resolve any disputes at its sole discretion. Providing false or misleading information including, but not limited to, identity, residency, contact details, or ownership of rights, will result in immediate disqualification. The Sponsor may request proof of residency or identity at any stage of the Competition to ensure compliance with rules mentioned above.
  • Data access and use: The provided VHH sequence data and evaluation data will be "accessible" (e.g., via streaming or loading) directly into the participants’ local environments or the provided computational environments. Ground Truth data will be kept private until the end of the submission evaluations.
  • External Data Restrictions: The use of external data is strictly limited to "publicly available datasets" (e.g., PDB, SAbDab). The use of proprietary or private data is prohibited, as it severely compromises scientific reproducibility. Participants using external data are obligated to clearly state its source and license.
  • Modeling Constraints and Reproducibility Obligations: There are no restrictions on the computational approaches (e.g., specific Deep Learning methods) used. For Track A, participants will be asked to provide sufficient documentation to support reproducibility of their submitted predictions, including model inputs, inference settings, and any public external data used. Track B will follow the standard CAPRI participation, submission, and assessment rules. Code submission will not be required for Track B.
  • Open Source Obligation: We will encourage top-performing Track A teams to share their inference workflows and code, where possible, to promote transparency and reproducibility. Any code release will respect participants’ intellectual property and applicable institutional or company policies.
  • Compulsory Licensing: Participants grant the SEPIQ organizers permission to use submitted Track A predictions for evaluation, leaderboard reporting, meta-evaluation, and academic publication of aggregate competition results. Track B data use and reporting will follow the standard CAPRI policies.
  • Communication Rules: Organizers are available for 24/7 support via support email. In case of rules or deadline changes, organizers will communicate them on the SEPIQ Competition Page, website, and LinkedIn page.

Competition promotion and incentives

The SEPIQ challenge targets researchers at different levels and industries: university students and faculty, industry researchers in various fields. We engage people who are interested in (but not limited to) computational biology, bioinformatics, drug discovery, AI and ML. Publication Plan and Co-authorship Invitations: Following the competition, the SEPIQ management team will lead a "meta-evaluation" (cross-analyzing multiple submissions and results to extract new insights in AI drug discovery and AI immunology). Top-ranking teams from each track will be invited as Consortium Authors on the summary paper detailing this meta-evaluation. We will prepare a summary manuscript for submission to an appropriate venue in machine learning or computational biology. Prize Distribution Structure: The confirmed prize amounts for Track A are: 1st Place $3,000, 2nd Place $2,000, and 3rd Place $1,000. These awards are intended to recognize top-performing participants while emphasizing the scientific value and academic visibility of the challenge. Any recognition or prize structure for the associated Track B will be determined separately in coordination with CAPRI. Diversity and Inclusion: To promote diversity, we will reserve travel awards specifically for participants from underrepresented groups and promote our challenge among organizations like Black in AI, LatinX in AI, Women in Machine Learning (WiML), and Queer in AI.

Resources

Organizing team

SEPIQ is organized by a very diverse international team of researchers, data scientists, academics, and industry partners. Responsible for the SEPIQ platform, challenge organization, proposal development, baseline model, starter kit creation, and competition promotion: Ihor Neporozhnii, Oleksandra Ostapenko, Hiromitsu Harimoto, Anastasiia Ryzhkova, Dr. Cheng-Hao Liu, Dr. Kentaro Tomii, Teruki Honma. COGNANO Inc. operation and training data production: Dr. Hiroyuki Yamazaki, Ryota Maeda, Dr. Akihiro Imura. Structural ground truth Cryo-EM complexes production: Dr. Keiichi Namba, Kasai Kazuki, Dr. Tsuyoshi Inoue. Dr. Takashi Nagata and Dr. Kliment Olechnoviˇc advise on evaluation methodology, including interface annotation, residue-level epitope metrics, structural-bioinformatics consistency, and interpretation of Track A and Track B evaluation results. Website operation: Tomohisa Oda.

Resources provided by organizers

SEPIQ organizers provide the following resources: 1) Starter Kit, which includes Jupyter Notebook with the baseline model and tutorial on dataset access; 2) Support Team, which is available for inquiries via info@cognano.co.jp; 3) monetary awards and, subject to sponsorship, limited travel support based on competition outcome.

References

  • The OpenFold3 Team. Openfold3-preview, 2025. URL https://github.com/aqlaboratory/openfold-3.
  • Rui Yin and Brian G. Pierce. Evaluation of alphafold antibody–antigen modeling with implications for improving predictive accuracy. Protein Science, 33(1):e4865, 2024. doi: https://doi.org/10.1002/pro.4865. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/pro.4865.
  • Francis Gaudreault, Traian Sulea, and Christopher R Corbeil. Ai-augmented physics-based docking for antibody–antigen complex prediction. Bioinformatics, 41(4):btaf129, 04 2025. ISSN 1367-4811. doi: 10.1093/bioinformatics/btaf129. URL https://doi.org/10.1093/bioinformatics/btaf129.
  • Fatima N. Hitawala and Jeffrey J. Gray. What does alphafold3 learn about antibody and nanobody docking, and what remains unsolved? mAbs, 17(1):2545601, 2025. doi: 10.1080/19420862.2025.2545601. URL https://doi.org/10.1080/19420862.2025.2545601. PMID: 40814020.
  • Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, {Andrew J.} Ballard, Joshua Bambrick, {Sebastian W.} Bodenstein, {David A.} Evans, {Chia Chun} Hung, Michael O’Neill, David Reiman, Kathryn Tunyasuvunakool, Zachary Wu, Akvil˙e Žemgulyt˙e, Eirini Arvaniti, Charles Beattie, Ottavia Bertolli, Alex Bridgland, Alexey Cherepanov, Miles Congreve, {Alexander I.} Cowen-Rivers, Andrew Cowie, Michael Figurnov, {Fabian B.} Fuchs, Hannah Gladman, Rishub Jain, {Yousuf A.} Khan, {Caroline M.R.} Low, Kuba Perlin, Anna Potapenko, Pascal Savy, Sukhdeep Singh, Adrian Stecula, Ashok Thillaisundaram, Catherine Tong, Sergei Yakneen, {Ellen D.} Zhong, Michal Zielinski, Augustin Žídek, Victor Bapst, Pushmeet Kohli, Max Jaderberg, Demis Hassabis, and {John M.} Jumper. Accurate structure prediction of biomolecular interactions with alphafold 3. Nature, 630(8016):493–500, June 2024. ISSN 0028-0836. doi: 10.1038/s41586-024-07487-w. Publisher Copyright: © The Author(s) 2024.
  • Chai Discovery, Jacques Boitreaud, Jack Dent, Matthew McPartlon, Joshua Meier, Vinicius Reis, Alex Rogozhnikov, and Kevin Wu. Chai-1: Decoding the molecular interactions of life. bioRxiv, 2024. doi: 10.1101/2024.10.10.615955. URL https://www.biorxiv.org/content/early/2024/10/15/2024.10.10.615955.
  • John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. Applying and improving alphafold at casp14. Proteins: Structure, Function, and Bioinformatics, 89(12):1711–1721, 2021. doi: https://doi.org/10.1002/prot.26257. URL https://onlinelibrary.wiley.com/doi/abs/10.1002/prot.26257.
  • Xiaoqiang Huang, Jun Zhou, Shuang Chen, Xiaofeng Xia, Y. Eugene Chen, and Jie Xu. Saaint-db: a comprehensive structural antibody database for antibody modeling and design. Acta Pharmacologica Sinica, 46(12):3365–3375, June 2025. doi: 10.1038/s41401-025-01608-5. URL https://www.nature.com/articles/s41401-025-01608-5#citeas.
  • Constantin Schneider, Matthew I J Raybould, and Charlotte M Deane. Sabdab in the age of biotherapeutics: updates including sabdab-nano, the nanobody structure tracker. Nucleic Acids Research, 50(D1):D1368–D1372, 01 2022. ISSN 0305-1048. doi: 10.1093/nar/gkab1050. URL https://doi.org/10.1093/nar/gkab1050.
  • James Dunbar, Konrad Krawczyk, Jinwoo Leem, Terry Baker, Angelika Fuchs, Guy Georges, Jiye Shi, and Charlotte M. Deane. Sabdab: the structural antibody database. Nucleic Acids Research, 42(D1):D1140–D1146, 01 2014. ISSN 0305-1048. doi: 10.1093/nar/gkt1043. URL https://doi.org/10.1093/nar/gkt1043.
  • Kilment Olechnovič, Sergei Grudinin. Voronota-LT: Efficient, Flexible, and Solvent-Aware Tessellation-Based Analysis of Atomic Interactions. J Comput Chem. 2025 Jul 15;46(19):e70178. doi: 10.1002/jcc.70178.
  • Sankar Basu and Björn Wallner. Dockq: A quality measure for protein-protein docking models. PLoS ONE, 11(8):e0161879, August 2016. doi: 10.1371/journal.pone.0161879. URL https://doi.org/10.1371/journal.pone.0161879.
  • KeeleyW. Collins, Matthew M. Copeland, Guillaume Brysbaert, Shoshana J.Wodak, Alexandre M.J.J. Bonvin, Petras J. Kundrotas, Ilya A. Vakser, and Marc F. Lensink. Capri-q: The capri resource evaluating the quality of predicted structures of protein complexes. Journal of Molecular Biology, 436(17):168540, 2024. ISSN 0022-2836. doi: https://doi.org/10.1016/j.jmb.2024.168540. URL https://www.sciencedirect.com/science/article/pii/S0022283624001359. Computation Resources for Molecular Biology.
  • Hirofumi Tsuruta, Hiroyuki Yamazaki, Ryota Maeda, Ryotaro Tamura, Jennifer N. Wei, Zelda E Mariet, Poomarin Phloyphisut, Hidetoshi Shimokawa, Joseph R. Ledsam, Lucy J Colwell, and Akihiro Imura. AVIDa-hIL6: A large-scale VHH dataset produced from an immunized alpaca for predicting antigen-antibody interactions. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. URL https://openreview.net/forum?id=BgY17iEnTb.
  • Justyna Bazan, Ireneusz Całkosi´nski, and Andrzej Gamian. Phage display—a powerful ntechnique for immunotherapy. Human Vaccines Immunotherapeutics, 8(12):1817–1828, November 2012. doi:10.4161/hv.21703. URL https://doi.org/10.4161/hv.21703.
  • George P. Smith. Filamentous fusion phage: novel expression vectors that display cloned antigens on the virion surface. Science, 228(4705):1315–1317, June 1985. doi: 10.1126/science.4001944. URL https://doi.org/10.1126/science.4001944.
  • Kaiwen Zhang, Yuhao Tao, and Fei Wang. Antibinder: utilizing bidirectional attention and hybrid encoding for precise antibody-antigen interaction prediction. Briefings in Bioinformatics, 26(1): bbaf008, 01 2025. ISSN 1477-4054. doi: 10.1093/bib/bbaf008. URL https://doi.org/10.1093/bib/bbaf008.

A Biography of all team members

  • SEPIQ team responsible for SEPIQ platform, challenge organization, baseline model, and starting kits creation, and competition promotion:
    • Ihor Neporozhnii is a PhD candidate and Research Assistant at the University of Toronto. His research is focused on accelerating the discovery of novel materials and medicines using machine learning and quantum chemistry.
    • Oleksandra Ostapenko is a Data Scientist at COGNANO Inc., holding an MSc. in Physics from the University of British Columbia. Her work interests lie at the intersection of science, data, and real-world applications.
    • Hiromitsu Harimoto is a PhD student at the Cambridge Stem Cell Institute, specializing in stem cell biology. His research focuses on understanding hematopoietic stem cells and their differentiation patterns, with an emphasis on how bone marrow niches influence cell fate.
    • Anastasiia Ryzhkova is a clinical medical physicist with interests in dosimetry, biophysics, computational modeling, and radiation therapy. She obtained her MSc. in Physics from the Taras Shevchenko National University of Kyiv.
    • Dr. Kentaro Tomii is a Professor at Artificial Intelligence Research Center, National Institute of Advanced Industrial Science and Technology.
    • Teruki Honma is a Team Director, RIKEN Center for Integrative Medical Science.
    • Dr. Cheng-Hao Liu is a postdoctoral fellow at Caltech, supervised by Prof. Frances Arnold specialising in chemistry and computer science.
  • COGNANO Inc. supports training data production:
    • Dr. Hiroyuki Yamazaki is the Chief of Hematology at Shizuoka City Shizuoka Hospital and a Data Scientist at COGNANO Inc. He aims to develop the fusion area between the life science field and the machine learning/IT fields.
    • Ryota Maeda is a Development Manager at COGNANO Inc., holding an MSc. in Life Science from Kyoto University. He makes large-scale labeled VHH datasets.
    • Dr. Akihiro Imura is a COGNANO Inc. President, CEO, and Co-founder. His research areas include cancer, homeostasis, antibodies, genetic engineering, and cell biology.
    • Tomohisa Oda is a Software Engineer at COGNANO Inc., started his career as a user interface designer and gained experience as a software engineer in a wide range of areas, from front-end to back-end design and development.
  • The University of Osaka Team provides ground truth structural Cryo-EM complexes:
    • Dr. Keiichi Namba is a Specially Appointed Professor at the Graduate School of Frontier Biosciences and JEOL YOKOGUSHI Research Alliance Laboratories, The University of Osaka. He specializes in structural biology and electron microscopy of biological macromolecular assemblies and bacterial flagellar motors. He supervised cryo-EM structural analysis of target complexes and provided resources.
    • Kazuki Kasai is a Specially Appointed Researcher at the Graduate School of Frontier Biosciences and JEOL YOKOGUSHI Research Alliance Laboratories, The University of Osaka. He specializes in structural biology using cryo-EM. He acquired cryo-EM data and performed image analysis to solve the structures of the target complexes.
    • Dr. Tsuyoshi Inoue is a Professor at the Graduate School of Pharmaceutical Sciences. His research interest is drug discovery with the theme of developing new modalities based on the structural biology method.
  • SEPIQ evaluation methodology team:
    • Dr. Takashi Nagata is an Associate Professor, Institute of Advanced Energy, Kyoto University. His research focuses on elucidating the functions and dysfunctions of proteins, DNA, RNA, and bioactive compounds associated with cancer, neurological disorders, and infectious diseases.
    • Dr. Kliment Olechnoviˇc is a Senior Researcher at Vilnius University Life Sciences Center, specializing in developing novel, effective methods of structural bioinformatics.

Acknowledgment

We thank Alexandre M.J.J. Bonvin and Marc Lensink for insightful discussions and the collaboration on the Track B CAPRI initiative. We also thank Stephen K. Burley, Mallory R. Tollefson, and Dima Kozakov for valuable discussions and the continuous support of the SEPIQ Competition. In addition, we thank Kaori Yurugi, Yasuko Imura, and Takatsugu Hirokawa for their assistance with organizing the SEPIQ Competition.