Consortium Launched to Build Antibody AI Dataset
The collaboration seeks to standardize antibody developability modeling by training AI on 10,000 distinct samples.
Updated on Sept. 29, 2026 in Biotech

Live Poll
Do you support pharmaceutical companies sharing private research data to improve AI drug development?
Apheris and Ginkgo Datapoints have launched the Antibody Developability Consortium to create a 10,000-antibody dataset for AI training. The research-stage initiative aims to overcome the limitations of small, fragmented data currently used in pharmaceutical development.
Why it matters
The consortium seeks to solve the lack of consistent, high-quality data that currently hinders the development of predictive models for antibody behavior. By aggregating proprietary sequences, the project aims to improve how drug candidates are screened and optimized for clinical success.
The project aims to compile a set of 10,000 antibodies to improve predictive performance over current fragmented datasets. This scale represents a significant attempt to build a unified benchmark for antibody developability endpoints.
The players
Apheris
A Berlin-based technology firm specializing in federated learning infrastructure for secure, privacy-preserving AI training.
Ginkgo Datapoints
A division of Ginkgo Bioworks focused on high-throughput laboratory automation and biological data generation.
AbbVie
A global research-based biopharmaceutical company with a extensive portfolio in immunology and oncology therapeutics.
Charlotte Deane
A professor and researcher providing independent scientific oversight for the project's data standards.
Peter Tessier
An academic expert providing independent oversight for the consortium's computational and biological modeling approaches.
The details
The consortium uses Apheris federated infrastructure—a system that allows AI models to train on decentralized data without moving the underlying sensitive sequences—to process proprietary data from members. Ginkgo Datapoints provides high-throughput wet-lab characterization to ensure data consistency. Scientists Charlotte Deane and Peter Tessier provide oversight to ensure the methodologies meet rigorous standards for drug discovery.
Timeline
September 29, 2026: The Antibody Developability Consortium officially launched.
Early 2027: The consortium plans to deliver the initial dataset to its members.
The Tech Race
The Antibody Developability Consortium marks a shift from isolated, firm-specific modeling to a collaborative data-sharing model. It directly attempts to build the industry-standard benchmark that has remained elusive due to the proprietary nature of pharmaceutical R&D.
For pharmaceutical researchers and developers, this effort signifies a potential improvement in the accuracy of early-stage drug screening tools. Industry members retain ownership of their sequences, meaning the platform functions as an infrastructure layer rather than a central database.
The takeaway
This effort represents a significant attempt to build a cross-industry standard for antibody modeling. Watch for the delivery of the initial dataset in early 2027 to see if the consortium successfully improves predictive outcomes compared to current methods.
What happens next
The consortium intends to provide its initial dataset to participating members in early 2027.
Further reading
For more on the latest research and industry trends, visit Biotech.
Source note: This article includes information reported by Tech.
Live Poll
Do you support pharmaceutical companies sharing private research data to improve AI drug development?







