ANTIBODY DEVELOPABILITY CONSORTIUM

Developability models you can trust on your own antibodies

A pre-competitive consortium building the industry’s largest standardized developability dataset: a 10,000-antibody panel characterized to a single set of protocols, with models delivered into your environment and your sequences never leaving it.

WHY THE CONSORTIUM EXISTS

Antibodies fail late, for reasons the data cannot yet predict.

Developability, whether a candidate can be manufactured, formulated and developed into a clinical product, decides which antibodies reach patients. Predicting it early would save years and significant investment, but the data needed to train those models has never existed at scale.

Datasets are siloed and non-standardized.

Models do not generalize.

No single company has enough.

The gap is measured, not assumed.

In a blinded benchmark run by Ginkgo Datapoints in 2025, 113 teams from 25 countries, 38 companies and 39 universities built developability predictors on 246 public clinical antibodies and were scored on 80 held-out molecules. The best submissions reached a Spearman ρ of 0.708 on the strongest property, but between 0.31 and 0.39 across the rest. The published conclusion was that available datasets are too small and too heterogeneous to support prediction that holds across assays.

2025 Ginkgo Datapoints Antibody Developability Competition outcomes, mAbs (2026)

THE SCIENCE BEHIND IT

Built on published, peer-reviewed method work.

THE PLATFORM

A high-throughput platform for biophysical antibody developability assessment to enable AI/ML model training

Arsiwala et al., mAbs (2025)

Documents the assay platform behind the consortium dataset and its run-to-run reproducibility across a 246-antibody benchmark.

Read the paper

THE BENCHMARK

2025 Ginkgo Datapoints Antibody Developability Competition outcomes

mAbs (2026)

113 teams scored against 80 blinded antibodies. Established that models trained on today’s public data do not generalize across assays.

Read the paper

THE OPEN DATA

GDPa1–GDPa4 developability datasets

datapoints.ginkgo.bio

246 IgGs, 18 VHH constructs, 80 further IgGs, and 160 bispecifics with 71 parental precursors, all publicly released.

Browse the datasets

HOW IT WORKS

Ginkgo builds the data. Apheris delivers the models.

The work splits cleanly in two. Ginkgo Datapoints generates the dataset and trains the foundation model; Apheris delivers that model into each member’s own environment and keeps every member’s sequences private.

Generates the data, trains the model

Selects sequences with a diversity algorithm built collectively with the pharma members and academic advisors to maximize coverage

Produces and purifies the IgGs

Runs high-throughput characterization across the core assay panel, screening 2,400+ antibody designs in parallel

Trains the foundation model on the resulting dataset

Delivers the models, protects the sequences

Delivers the foundation model into each member’s own environment

Members fine-tune on proprietary data, or train new models on consortium data

Runs the diversity algorithm locally, so sequence selection happens inside each member’s environment

Protects cross-member privacy of contributed sequences in the central dataset

WHAT GETS BUILT

10,000 antibodies, characterized to one standard.

Ginkgo Datapoints selects 10,000 diverse IgG sequences and characterizes them to a single set of protocols, beginning with the first 5,000 in Phase 2. Members contribute proprietary sequences; Ginkgo fills the remainder from public sources.

Standardized characterization

Every antibody in the set is characterized on the same platform, to a single set of protocols, so all members work from consistent, comparable data. The property panel is shared with prospective members in the consortium overview.

Two phases to the first model

Phase 1 · Sequence selection

1–2 MONTHS

10,000 diverse IgG sequences selected. Ginkgo’s diversity algorithm runs locally inside each member’s environment.

Phase 2 · Data generation + model training

4–5 MONTHS

The first 5,000 IgGs produced and characterized. A foundation model is trained and benchmarked.

Beyond Phase 2, the consortium may extend into further characterization or additional properties such as viscosity, depending on how the Phase 2 model performs.

WHAT IT CHANGES

Fewer re-engineering loops, and earlier decisions to stop.

A developability liability found after expression costs months. Prediction that holds up prospectively moves that decision to the design stage. The consortium model is where that starts, not where it ends.

Run it where your data is

The consortium model is delivered through Apheris Foundry into your own infrastructure. Nothing has to leave it.

Fine-tune per drug program

Specialize the model on one program’s sequences and assay history, where the liabilities are specific and a general model is weakest.

Train your own models

Federated access lets you train new architectures against every contributed sequence, not only run the model the consortium trained.

Keep it current

Each phase adds characterized antibodies, and your programs keep generating more. Retrain, and predictions move with the portfolio.

WHAT MEMBERS GET

What each member receives.

Ginkgo characterizes every antibody in the set to a single standard, so all members work from consistent, comparable data.

The consortium model

A trained foundation model, weights plus reproducible code, for predicting key developability properties.

Federated access to the full dataset

Train, benchmark and fine-tune your own models on the full dataset, without data leaving a secure environment.

Your raw assay data

Raw assay data for every sequence you contribute, plus the assay data generated from the publicly sourced sequences.

Steering input

Phase-by-phase reports and joint technical meetings to review results and steer the direction of the project.

You keep what you bring.

You retain full ownership of the sequences you contribute and the assay data generated from them.

Whatever you build from the consortium, models trained on the data, or fine-tuned from the consortium model, is yours to use internally, without restriction.

FAQ

Frequently asked questions about AbDev.

FIRST CONVERSATION

What happens when you reach out.

José-Tomás (JT) Prieto

Director of AI Programs
— business and legal

Leonardo Castorina

Senior ML Engineer, Large Molecules — the science

Your first conversation is with José-Tomás and Leonardo. These are your Apheris contacts; Ginkgo Datapoints leads the data generation and joins for the scientific design.

  1. Understand your portfolio.
    We talk through your antibody portfolio, which sequences you could contribute, and where developability decisions are costing you today.
  2. Check the scientific fit.
    Leo walks through the dataset, the assay panel and the modelling plan, and what is realistic to expect on your own molecules.
  3. Membership and terms.
    JT covers how membership works, the sequence-handling and IP arrangements, and what it would cost. Your IP team can be involved from the start.

JOIN OUR NETWORK

Join the Antibody Developability Consortium

Sequence selection happens once per phase, so members who join earlier get more of their own molecules into the dataset and into the model trained on it. Places are limited, and enrollment closes once data generation is underway.

A first conversation covers your sequence portfolio, how membership would work, and what it would cost.

The consortium is jointly led by Ginkgo Datapoints and Apheris. Either team can take you through it.