
Artificial intelligence has become increasingly important in drug discovery, but one of its biggest advantages is also one of the industry’s biggest problems: data.
The most useful molecular information is often not sitting in a public database. It is inside pharmaceutical companies, generated during years of drug-discovery programmes and protected as valuable intellectual property.
That creates a difficult trade-off. AI models need large and diverse datasets to improve, but pharmaceutical companies have little incentive — and often strong commercial reasons not — to hand their proprietary molecular structures to competitors.
A new collaboration involving five pharmaceutical companies and researchers at Columbia University has demonstrated a possible way around that problem.
AbbVie, Astex Pharmaceuticals, Bristol Myers Squibb, Johnson & Johnson and Takeda jointly fine-tuned an AI model using more than 20,000 proprietary protein–ligand structures. The crucial detail is that the companies did not pool those structures in a common database.
Instead, they used a technique known as federated learning. Each company trained the model inside its own computing environment, while only model updates were aggregated between participants.
The resulting model, called AISB-1-Fed, showed a substantial improvement on a held-out set of 1,056 private structures. The consortium reported that high-quality protein–ligand interface predictions increased from 35.6% for the public OpenFold3 Preview 2 model to 52.1% for AISB-1-Fed. Correct ligand-pose predictions also increased from 28.9% to 46.8%.
The results were reported by the AI Structural Biology Network in September 2026 and covered independently by Nature. The work has not yet been peer-reviewed, and the resulting model remains private.
Why is pharmaceutical data so valuable for AI drug discovery?
AI models for biology have benefited enormously from public structural databases.
The Protein Data Bank, for example, contains hundreds of thousands of experimentally determined protein structures and has been an important resource for computational biology. But the public record contains relatively few examples of proteins interacting with drug-like small molecules.
That distinction matters.
Knowing the three-dimensional structure of a protein is different from knowing how that protein interacts with a potential drug molecule.
Drug discovery often depends on understanding precisely how a small molecule fits into a protein’s binding site. Researchers can then use that information to design or modify compounds intended to interact with the target.
Nature reported that the public Protein Data Bank contains perhaps 10,000 structures involving proteins bound to drug-like molecules, according to Astex’s Paul Mortenson. Pharmaceutical companies, meanwhile, possess substantial collections of structures generated during proprietary drug-development programmes.
Those private datasets can contain exactly the kind of information that public AI models lack.
The problem is that companies are unlikely to simply upload their internal structures into a shared repository.
A structure can reveal information about a company’s research programme, its targets, chemical series and potential drug candidates. Sharing such information could therefore expose commercially sensitive research.
Federated learning offers another route.
What is federated learning?
Federated learning changes where the training happens.
In conventional machine learning, researchers generally assemble the training data in one location. The model is then trained against that combined dataset.
Federated learning reverses that arrangement.
Each participating organisation keeps its data inside its own environment. A copy of the model is sent to the organisation, where it is trained against local data. Instead of sending the underlying data back, the organisation contributes model updates or parameters.
Those updates can then be aggregated to improve a shared model.
The process can be repeated across participating organisations.
In the AISB collaboration, the pharmaceutical companies therefore did not need to create a central database containing their proprietary protein–ligand structures. Each company trained OpenFold3 locally on its own structures, while the model parameters were aggregated through the federated infrastructure.
The result was effectively a model that could learn from a much larger combined body of evidence without requiring the underlying molecular structures to be exchanged.
The distinction is important
Federated learning does not mean that the participating companies suddenly made their proprietary datasets public.
The structures remained within the companies’ environments. The trained AISB-1-Fed model and its underlying weights are also not publicly available. What has been disclosed publicly are the aggregate results of the evaluation.
That makes the project different from simply creating a shared pharmaceutical database.
What did the five companies actually train?
The collaboration built on OpenFold3 Preview 2, an open-source model designed to predict structures involving proteins and other molecules.
The participating companies fine-tuned the model using 20,167 experimentally determined protein–ligand structures from their drug-discovery programmes. The collaboration was conducted with the AlQuraishi Lab at Columbia University, which developed OpenFold3.
The five participating companies were:
- AbbVie
- Astex Pharmaceuticals
- Bristol Myers Squibb
- Johnson & Johnson
- Takeda
The data were kept within the respective companies’ environments during training.
According to Apheris, which provided the federated computing infrastructure, the training involved more than 20,000 proprietary structures and was completed in under 10 weeks.
That is significant not simply because of the number of structures, but because of their character.
The data came from real pharmaceutical discovery programmes rather than being assembled solely from public structural repositories.
The model showed a substantial jump in prediction quality
The consortium evaluated AISB-1-Fed on 1,056 structures that had been held out from training.
The results showed a sizeable difference between the public starting model and the federated model.
High-quality protein–ligand interface predictions
AISB-1-Fed achieved the required high-quality interface prediction on 52.1% of the held-out structures.
OpenFold3 Preview 2 achieved the same threshold on 35.6%.
For comparison, the public Boltz-2 model reached 40.9% on the same evaluation, according to the consortium’s results.
That means AISB-1-Fed improved the starting model’s result by 16.5 percentage points.
Correct ligand poses
The second measure examined whether the model correctly positioned the drug-like molecule relative to its protein target.
AISB-1-Fed achieved a correct pose on 46.8% of the evaluated structures.
The starting OpenFold3 model reached 28.9%, while Boltz-2 reached 36.5%.
The improvement over the starting model was therefore 17.9 percentage points.
These are benchmark improvements rather than evidence that a particular drug candidate will succeed in clinical development.
That distinction matters.
What does this mean for drug discovery?
The immediate implication is not that AI has suddenly solved drug discovery.
Rather, the experiment addresses a specific bottleneck: how to train models on valuable proprietary scientific data when companies cannot or will not centralise that data.
Better structural predictions can potentially help researchers understand protein–drug interactions earlier in the discovery process.
If a model can more accurately predict how a candidate molecule may sit inside a protein’s binding site, researchers can use those predictions when evaluating and designing compounds.
That could help reduce some unnecessary experimental work.
But there are many stages between a computational prediction and an approved medicine.
A model still has to contend with questions involving potency, selectivity, toxicity, pharmacokinetics, formulation, manufacturing, animal studies and clinical development.
AISB-1-Fed itself is focused on structural prediction. It is not a system that directly predicts whether a drug will work in humans.
The consortium’s next efforts are moving toward models that can predict how tightly small molecules bind to their targets, which would address a different problem from structural co-folding.
Why the federated approach could matter beyond pharma
The larger idea extends beyond molecular biology.
Many industries possess valuable datasets that cannot easily be pooled.
Hospitals have patient records. Banks have sensitive financial information. Manufacturers have proprietary production data. Energy companies have operational measurements. Governments hold restricted datasets.
In each case, there can be a tension between data utility and data control.
A conventional AI project often works best when large quantities of data can be brought together. Privacy rules, intellectual-property concerns, competition issues or commercial incentives can make that difficult.
Federated learning provides another model: allow the algorithm to move to the data rather than moving the data to the algorithm.
The pharmaceutical experiment therefore demonstrates a potentially important principle.
Competitors do not necessarily have to surrender their underlying datasets to benefit from collective machine learning.
That does not eliminate the legal, technical or commercial challenges of collaboration. It does, however, change the range of possible architectures for handling sensitive data.
The biggest caveat: this is not yet independent proof
The results are notable, but they need to be interpreted within the limits of the experiment.
The findings were published by the consortium through Apheris and covered by Nature. The work had not been peer-reviewed when Nature reported it, and the consortium said it planned to submit a paper for peer review.
The evaluation was also conducted on structures held out by the participating companies themselves.
That makes the benchmark useful, but it is not the same as demonstrating performance on an entirely independent external dataset.
The public does not have access to the underlying structures or the trained AISB-1-Fed weights either. Consequently, independent researchers cannot currently reproduce the full experiment themselves.
There is another important limitation.
The reported improvement concerns protein–ligand structural prediction. It should not be described as an improvement in binding affinity prediction.
A model that predicts where a molecule sits in a protein’s binding pocket is solving a related but distinct problem from predicting how strongly the molecule binds.
That distinction becomes particularly important when discussing potential applications in real-world drug development.
What happens next?
The next test for the approach will be whether the reported gains survive broader and more independent evaluation.
Peer review will provide one important check.
External benchmarks could provide another. Independent researchers would also benefit from greater transparency around evaluation datasets, target classes and model performance.
At the same time, the collaboration is expanding its ambitions.
The AISB Network has launched another federated initiative aimed at predicting how tightly small molecules bind to protein targets, with AbbVie, AstraZeneca, Bristol Myers Squibb and Johnson & Johnson participating. The stated applications include virtual screening and lead optimization.
That development points toward a broader experiment: whether federated learning can become a continuing infrastructure for pharmaceutical AI rather than a one-time demonstration.
If it can, companies could potentially collaborate on increasingly sophisticated models while retaining control over the underlying datasets that give them a competitive advantage.
The real breakthrough may be the architecture, not the model
The headline number is the jump from 35.6% to 52.1%.
But the more consequential development may be what made that jump possible.
Five pharmaceutical competitors contributed knowledge to a common AI training process without creating a shared repository of their proprietary molecular structures.
That is a fundamentally different way of thinking about collaboration.
For years, the basic problem in AI has often been framed as a shortage of data. In some scientific fields, however, the data exists in abundance — it is simply distributed across organisations that cannot realistically combine it.
Federated learning offers a way to work around that constraint.
AISB-1-Fed does not prove that every sensitive-data AI project can benefit from federation. It does not eliminate the need for security, governance, contractual arrangements or independent validation.
But the experiment provides a concrete example of something that previously looked difficult for obvious commercial reasons: rival pharmaceutical companies improving a common scientific model without handing each other their molecular data.
The next question is whether that idea can move from an impressive benchmark to a reproducible tool that changes how medicines are actually discovered.
For now, the answer remains open.
TL;DR
- AbbVie, Astex Pharmaceuticals, Bristol Myers Squibb, Johnson & Johnson and Takeda jointly fine-tuned OpenFold3 using 20,167 proprietary protein–ligand structures.
- The structures remained inside the participating companies’ environments through federated learning.
- AISB-1-Fed achieved high-quality interface predictions on 52.1% of 1,056 held-out structures, compared with 35.6% for the public OpenFold3 Preview 2 baseline.
- Correct ligand poses increased from 28.9% to 46.8%.
- The model also outperformed the public Boltz-2 benchmark used in the comparison.
- The work has not yet been peer-reviewed, and the model and underlying private datasets are not publicly available.
- The experiment demonstrates a method for collaborative AI training on commercially sensitive scientific data; it does not by itself demonstrate faster clinical development or a successful new drug.