Pharma Firms Trained AI Without Sharing Proprietary Data

A collaborative project utilized over 20,000 secret molecular structures to outperform existing public AI models.

Updated on Oct. 1, 2026 in Artificial Intelligence

Isometric editorial illustration of a complex molecular structure made of geometric spheres and rods, representing collaborative artificial intelligence research.
Five pharmaceutical firms successfully developed an AI model to predict drug interactions by sharing private model updates rather than raw molecular data. AI Illustration. Upload story photo >

Live Poll

Should competing companies collaborate to train shared AI models without pooling their private data?

Five pharmaceutical companies successfully trained an AI model on sensitive molecular data without pooling their confidential information. The project, completed in under ten weeks, demonstrated that collaborative learning can effectively predict protein and drug molecule interactions.

Why it matters

Pharmaceutical companies traditionally guard molecular structures as trade secrets, preventing the massive data pooling usually required for high-performance AI. This method bypasses that barrier, potentially accelerating drug discovery across the entire industry.

The jointly trained AISB-1-Fed model predicted protein and drug molecule interactions for 52.1% of test structures, significantly outpacing the 35.6% accuracy of the public model benchmark.

The players

Columbia University

An academic institution with a research laboratory that provided the collaborative framework for this AI training project.

AbbVie

A research-based biopharmaceutical company that collaborated on the AI model training using proprietary molecular data.

Bristol Myers Squibb

A global biopharmaceutical firm that contributed to the private data pooling initiative for predictive drug modeling.

Johnson & Johnson

A multinational corporation focused on pharmaceuticals and medical technology that participated in the federated training study.

Takeda

A research-driven pharmaceutical company that utilized its private molecular dataset to help train the collaborative AISB-1-Fed model.

The details

The team utilized a federated learning approach, where each company trained a local copy of the model on its own private data, then shared only the model updates rather than the raw information. This allowed the system to learn from 20,000 proprietary molecular structures without compromising the confidentiality of the companies involved. The final model effectively integrated these varied datasets to improve predictive accuracy for complex biochemical interactions.

Timeline

  1. October 2026: Publication of the collaborative project results.

The Tech Race

This project applies the federated learning paradigm to the high-stakes, data-siloed domain of drug discovery. It marks a departure from conventional centralized model training, establishing a new milestone for secure cooperation between commercial competitors.

This development could eventually lead to faster drug discovery cycles as the new training methodology is adopted by broader industry segments. Researchers and developers should monitor for upcoming peer-reviewed publications that will detail the model's reliability for clinical applications.

The takeaway

The success of this collaboration suggests that industry-wide data barriers can be overcome through privacy-preserving machine learning. Observers should watch for the forthcoming peer review of the AISB-1-Fed model to confirm its performance metrics.

Further reading

For more on the latest research in this field, visit Artificial Intelligence.

Source note: This article includes information reported by WION.

Live Poll

Should competing companies collaborate to train shared AI models without pooling their private data?