ANTIBODY DEVELOPABILITY CONSORTIUM
Developability models you can trust on your own antibodies
A pre-competitive consortium building the industry’s largest standardized developability dataset: a 10,000-antibody panel characterized to a single set of protocols, with models delivered into your environment and your sequences never leaving it.

WHY THE CONSORTIUM EXISTS
Antibodies fail late, for reasons the data cannot yet predict.
Developability, whether a candidate can be manufactured, formulated and developed into a clinical product, decides which antibodies reach patients. Predicting it early would save years and significant investment, but the data needed to train those models has never existed at scale.
Datasets are siloed and non-standardized.
Models do not generalize.
No single company has enough.
The gap is measured, not assumed.
In a blinded benchmark run by Ginkgo Datapoints in 2025, 113 teams from 25 countries, 38 companies and 39 universities built developability predictors on 246 public clinical antibodies and were scored on 80 held-out molecules. The best submissions reached a Spearman ρ of 0.708 on the strongest property, but between 0.31 and 0.39 across the rest. The published conclusion was that available datasets are too small and too heterogeneous to support prediction that holds across assays.
2025 Ginkgo Datapoints Antibody Developability Competition outcomes, mAbs (2026)
THE SCIENCE BEHIND IT
Built on published, peer-reviewed method work.
THE PLATFORM
A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training
Arsiwala et al., mAbs (2025)
Documents the assay platform behind the consortium dataset and its run-to-run reproducibility across a 246-antibody benchmark.
THE BENCHMARK
2025 Ginkgo Datapoints Antibody Developability Competition outcomes
mAbs (2026)
113 teams scored against 80 blinded antibodies. Established that models trained on today’s public data do not generalize across assays.
THE OPEN DATA
GDPa1–GDPa4 developability datasets
datapoints.ginkgo.bio
246 IgGs, 18 VHH constructs, 80 further IgGs, and 160 bispecifics with 71 parental precursors, all publicly released.
HOW IT WORKS
Ginkgo builds the data. Apheris delivers the models.
The work splits cleanly in two. Ginkgo Datapoints generates the dataset and trains the foundation model; Apheris delivers that model into each member’s own environment and keeps every member’s sequences private.
Generates the data, trains the model
Selects sequences with a diversity algorithm built collectively with the pharma members and academic advisors to maximize coverage
Produces and purifies the IgGs
Runs high-throughput characterization across the core assay panel, screening 2,400+ antibody designs in parallel
Trains the foundation model on the resulting dataset
Delivers the models, protects the sequences
Delivers the foundation model into each member’s own environment
Members fine-tune on proprietary data, or train new models on consortium data
Runs the diversity algorithm locally, so sequence selection happens inside each member’s environment
Protects cross-member privacy of contributed sequences in the central dataset
WHAT GETS BUILT
10,000 antibodies, characterized to one standard.
Ginkgo Datapoints selects 10,000 diverse IgG sequences and characterizes them to a single set of protocols, beginning with the first 5,000 in Phase 2. Members contribute proprietary sequences; Ginkgo fills the remainder from public sources.
Standardized characterization
Every antibody in the set is characterized on the same platform, to a single set of protocols, so all members work from consistent, comparable data. The property panel is shared with prospective members in the consortium overview.
Two phases to the first model
Phase 1 · Sequence selection
1–2 MONTHS
10,000 diverse IgG sequences selected. Ginkgo’s diversity algorithm runs locally inside each member’s environment.
Phase 2 · Data generation + model training
4–5 MONTHS
The first 5,000 IgGs produced and characterized. A foundation model is trained and benchmarked.
Beyond Phase 2, the consortium may extend into further characterization or additional properties such as viscosity, depending on how the Phase 2 model performs.
WHAT IT CHANGES
Fewer re-engineering loops, and earlier decisions to stop.
A developability liability found after expression costs months. Prediction that holds up prospectively moves that decision to the design stage. The consortium model is where that starts, not where it ends.
Run it where your data is
The consortium model is delivered through Apheris Foundry into your own infrastructure. Nothing has to leave it.
Fine-tune per drug program
Specialize the model on one program’s sequences and assay history, where the liabilities are specific and a general model is weakest.
Train your own models
Federated access lets you train new architectures against every contributed sequence, not only run the model the consortium trained.
Keep it current
Each phase adds characterized antibodies, and your programs keep generating more. Retrain, and predictions move with the portfolio.
Apheris brings considerable expertise from other data networks, including the ADMET Network and the AI Structural Biology Network.
WHAT MEMBERS GET
What each member receives.
Ginkgo characterizes every antibody in the set to a single standard, so all members work from consistent, comparable data.
The consortium model
A trained foundation model, weights plus reproducible code, for predicting key developability properties.
Federated access to the full dataset
Train, benchmark and fine-tune your own models on the full dataset, without data leaving a secure environment.
Your raw assay data
Raw assay data for every sequence you contribute, plus the assay data generated from the publicly sourced sequences.
Steering input
Phase-by-phase reports and joint technical meetings to review results and steer the direction of the project.
You keep what you bring.
You retain full ownership of the sequences you contribute and the assay data generated from them.
Whatever you build from the consortium, models trained on the data, or fine-tuned from the consortium model, is yours to use internally, without restriction.
FAQ
Frequently asked questions about AbDev.
FIRST CONVERSATION
What happens when you reach out.

José-Tomás (JT) Prieto
Director of AI Programs
— business and legal

Leonardo Castorina
Senior ML Engineer, Large Molecules — the science
Your first conversation is with José-Tomás and Leonardo. These are your Apheris contacts; Ginkgo Datapoints leads the data generation and joins for the scientific design.
JOIN OUR NETWORK
Join the Antibody Developability Consortium
Sequence selection happens once per phase, so members who join earlier get more of their own molecules into the dataset and into the model trained on it. Places are limited, and enrollment closes once data generation is underway.
A first conversation covers your sequence portfolio, how membership would work, and what it would cost.
The consortium is jointly led by Ginkgo Datapoints and Apheris. Either team can take you through it.