医疗领域独有数据增强了大模型——OpenFlod 3 , Drug firms’ secret data supercharge AI protein models

发布时间:2026/9/24 8:50:44

医疗领域独有数据增强了大模型——OpenFlod 3 , Drug firms’ secret data supercharge AI protein models 制药企业联盟AI Structural Biology | ApherisFederated Training Dramatically Improves the Accuracy of Protein-Ligand Co-folding on Private Pharma Structures五家制药公司在 20,167 个私有结构上对 OpenFold3 Preview 2 进行了微调这些数据均未离开各自的环境。在 1,056 个留出结构中高质量界面预测准确率从 35.6% 提升至 52.1%正确配体构象从 28.9% 提升至 46.8%超越了所有公开模型。这是联邦学习在药物发现领域的首次突破。ApherisFederated Training Dramatically Improves the Accuracy of Protein-Ligand Co-folding on Private Pharma StructuresFive pharma companies fine-tuned OpenFold3 Preview 2 across 20,167 private structures, none of which left their environments. On 1,056 held-out structures, high-quality interface predictions rose from 35.6% to 52.1% and correct ligand poses from 28.9% to 46.8%, ahead of every public model. A first for federated drug discovery.Public co-folding models are trained on public data, where only a small fraction of structures are drug-relevant, systematically under-representing the regime where structure-based drug design operates. The structures that carry the richest signal for drug-discovery models, dense medicinal-chemistry series around real targets, sit inside individual pharmaceutical companies, behind legal and commercial walls that keep them from being pooled.Federated learning offers a way around this. A shared model is trained across private datasets without the underlying structures ever leaving their owners. Additional privacy safeguards ensure no IP-sensitive information is exposed. Federation is not new to drug discovery. Earlier large-scale pharma federation efforts proved that cross-company training was feasible but delivered only modest gains in model performance. Making federation deliver a meaningful improvement is the problem Apheris set out to solve.Federated learning is the approach behind the AI Structural Biology (AISB) Network, an industry-led collaboration formed to advance AI for drug discovery, powered by Apheris.In their Federated OpenFold3 Initiative, participating AISB Network members asked whether OpenFold3 Preview 2, a publicly available co-folding model, fine-tuned jointly across their private structures, would predict protein-ligand complexes more accurately than public models, or than any model a single company could fine-tune on its own data alone.The answer is yes.The resulting federated model, which we call AISB-1-Fed, significantly raised the share of high-quality predictions from 36% to 52%, outperforming both public models and any model fine-tuned on a single partners data alone. To our knowledge, this is the first clear demonstration that federated training delivers a step-change in co-folding accuracy, the start of a new era for federated drug discovery.Run in collaboration with the AlQuraishi Lab at Columbia University, the developers of OpenFold3, AbbVie, Astex, Bristol Myers Squibb, Johnson Johnson, and Takeda each fine-tuned OpenFold3 Preview2 locally on their own structures, with model parameters periodically aggregated through Apheris’ federated computing product. In total, the training data spanned more than 20,000 experimentally determined structures from active drug-discovery programs. That roughly triples the drug-relevant protein-ligand data available for training, forming the most diverse such dataset assembled in drug discovery to date. By necessity, the underlying structures and the trained model weights remain private to the participating companies; what we are making public are the aggregate results.Main resultAISB-1-Fed substantially improves over OpenFold3 Preview 2 (OF3p2), the foundation model used for federation, and over all reference models evaluated on this private benchmark.AISB-1-Fed is OF3p2 after federated fine-tuning on the five partners roughly 20,000 private structures, together with public PDB data released up to November 2025. To separate the effect of the private data from that of the additional newer public data, the network also built a public-only baseline, AISB-1-OF3p2-all-PDB, by training OF3p2 on all available public PDB but no private structures. As the graphic shows, AISB-1-Fed clearly outperforms this baseline, so it is the private structures, that drive the improvement.Figure 1: Model comparison: fraction of structures with PL-lDDT ≥ 0.8 (ranked by confidence).We report protein-ligand interface lDDT (PL-lDDT) for local interface quality and ligand bisyRMSD for pose accuracy.AISB-1-Fed reaches 52.1% of structures with PL-lDDT ≥ 0.8, compared with 35.6% for OF3p2: a16.5 percentage pointgain (46% relative). For ligand pose accuracy, it reaches 46.8% with bisyRMSD ≤ 2 Å, compared with 28.9%: a17.9 percentage pointgain (62% relative).Against the strongest public reference model on this benchmark (Boltz-2), AISB-1-Fed leads by approximately11 percentage points on both metrics.‍Figure 2: Fraction of structures with PL-lDDT ≥ 0.8 versus fraction with bisyRMSD ≤ 2 Å. Up and to the right is better. The dashed arrow shows the improvement from OF3p2 to AISB-1-Fed .‍Figure 3: Fraction PL-lDDT ≥ 0.8 (left) and fraction bisyRMSD ≤ 2 Å (right) with bootstrapped 95% CIs (500 replicates, seed resampling). AISB-1-Fed leads on both metrics with non-overlapping CIs relative to all reference models.‍The two metrics are complementary.bisyRMSD focuses on absolute ligand placement, while PL-lDDT additionally captures interfacial interactions and pocket geometry. AISB-1-Fed improves both simultaneously, which makes the result more robust than an improvement on either metric alone.Evaluation setupWe evaluated models on private evaluation structures from five pharmaceutical partners. After keeping only evaluation structures with successful predictions and metric calculations for all compared models and checkpoints, this gives1,056 privateevaluationstructuresacross five datasets.Each partner held out 5% of its private structures for evaluation. The split was made at the project level, with all structures from a given medicinal-chemistry project assigned entirely to either training or evaluation. Partners were also asked to select evaluation projects that were unlikely to appear in any other partners training data. Together, these steps reduce the chance that evaluation structures are overly similar to training structures, whether from the same partner or from another. The five evaluation sets were pooled and treated as a single benchmark; no partner-level breakdown is shown without explicit partner approval.Training was run to a pre-specified compute budget, and the final checkpoint was reported.At a glance20,167 private structures for training and a 1,056 held-out set for validation from five pharmaceutical partnersTarget/project-level split within each partner; pooled across partners for reportingRanked sample selection using each models native confidence scorePrimary metrics: fraction PL-lDDT ≥ 0.8 and fraction bisyRMSD ≤ 2 ÅAggregate reporting only; no partner-level disclosure without approval‍‍Reference models.We compare against Boltz-2 and two ProtenixV1 checkpoints as external baselines, against OF3p2 as our starting-point reference, and against AISB-1-OF3p2-all-PDB as an internal control:OF3p2(technical report): the base model that AISB-1-Fed was fine-tuned from.AISB-1-OF3p2-all-PDB: OF3p2 fine-tuned within the AISB Network on public PDB only, with a later cutoff (2025-11-19) matching Boltz-2 and ProtenixV1-20250630. A conservative control isolating the effect of updated public data from the private pharmaceutical structures.Boltz-2(preprint): a co-folding model that jointly predicts complex structure and binding affinity.ProtenixV1(preprint): ByteDances co-folding model. We includeProtenixV1-Default(data cutoff aligned to AlphaFold3s original cutoff) andProtenixV1-20250630(extended cutoff for real-world use).‍All reference models were evaluated on the same 1,056 private structures with the same metric pipeline.Sample selection:Each structure was predicted with 5 diffusion samples across 5 seeds (25 predictions per structure). For each model, the best prediction is selected using the models own native ranking score, specifically the confidence heuristic it exposes for choosing among multiple generated structures. For OF3p2, AISB-1-Fed, and ProtenixV1, this is a composite score weighting interface confidence (ipTM), overall alignment confidence (pTM), and penalties for disorder and clashes. For Boltz-2, it is a weighted average of complex pLDDT and interface confidence.Thresholded fractions:PL-interface lDDT is bimodal on this benchmark. Roughly 28% of per-structure values fall below 0.2 (failures) and another 28% above 0.9 (near-perfect), with the middle largely empty. BisyRMSD has a long tail, with 29–41% of structures exceeding 12 Å depending on model. Means land in regions of the distribution that few structures actually inhabit. Thresholded fractions better answer the question of how often the model produces a useful prediction.Fraction PL-lDDT ≥ 0.8: a widely used marker of high-confidence interface qualityFraction bisyRMSD ≤ 2 Å: the conventional threshold for a correctly docked ligand pose‍Figure 4: Ridge KDE plots of per-structure PL-interface lDDT (left) and ligand bisyRMSD (right), ranked selection. Dashed lines mark the thresholds used in the headline metrics (PL-lDDT ≥ 0.8 is good; bisyRMSD ≤ 2 Å is good). For bisyRMSD the distribution is cli‍Why private pharmaceutical data helpsThe result is consistent with private pharmaceutical structures carrying a learning signal that public data does not fully cover. Public databases hold very few dense series of related compounds medicinal-chemistry series around real therapeutic targets. Since the original AlphaFold3 cutoff, the PDB has grown by roughly 70,000 structures, but almost all are cofactors, metabolites, ions, and crystallographic additives rather than drug-like compounds. Only about 10,600 public structures contain an approved or investigational drug, and only about 3,000 of those are new since that cutoff.The private structures contributed to AISB-1-Fed are all drug-discovery compounds, the very data the public record lacks.Protein-protein interface qualityThe Federated OpenFold3 Initiative was structured to improve protein-ligand co-folding; small molecule binding was the explicit target. We did not optimize for protein-protein interface (PPI) quality. Yet the PPI metric improved substantially as a by-product.Many protein-ligand complexes in this benchmark involve multimeric protein assemblies with biologically relevant protein-protein interfaces. For the443 structures with PPI annotations,we additionally report protein-protein interface lDDT (PPI-lDDT).‍ModelFrac PPI-lDDT ≥ 0.895% CIProtenixV1Default0.398[0.391, 0.404]ProtenixV1-202506300.391[0.386, 0.397]Boltz-20.409[0.406, 0.415]OF3p20.388[0.372, 0.400]AISB-1-OF3p2-all-PDB0.407[0.395, 0.418]AISB-1-Fed0.626[0.615, 0.634]AISB-1-Fed reaches 62.6% PPI-lDDT ≥ 0.8 on this subset, compared with 38.8–40.9% for the reference models, the largest aggregate shift in the benchmark, applying to the annotated PPI subset only.‍‍Figure 5: Fraction PPI-lDDT ≥ 0.8 on the 443-structure PPI subset, with 95% CIs.‍Figure 6: Ridge KDE of per-structure PPI-lDDT on the PPI subset. AISB-1-Fed has substantially more mass above 0.8 and a sharp peak near 1.0.‍How the AISB-1-Fed model was trainedThe AISB Network’s philosophy is to keep the federated algorithms simple, make the data interface consistent, and evaluate under realistic operational constraints. This approach has been validated in earlier federated experiments. It starts from OF3p2 and uses a 95/5 train/validation split of private structures. Training was accelerated using NVIDIA cuEquivariance kernels and distributed across multiple P5 instances powered by AWS EC2.Each participating organization prepared its local structural data in a shared OF3-compatible format. Federated learning sends the model to the data rather than the data to the model, so each partners structures stay in place while the shared model learns from all of them. Training used synchronous federated averaging: partners trained locally for a small number of steps, sent model parameters to a central aggregator (operated by Apheris using the Apheris Gateway), and received back the updated global model. Public PDB-derived data were mixed into training to stabilize optimization and prevent over-specialization to any one partners distribution, using a cutoff of2025-11-19.The private structures remained inside partner environments throughout. The aggregator received privacy-preserving model parameters, not structures.‍Figure 7: Federated learning workflow: local training inside each partner environment with central parameter aggregation.Data preparation and validationGetting the data ready was one of the biggest challenges of the project. Each company has built up its structures over decades in its own in-house formats and conventions, so making them consistent and machine-learning-ready across partners was far from automatic. Apheris and the AlQuraishi Lab built a shared>Artificial-intelligence-based models of protein structure could be improved by incorporating data from pharmaceutical companies.Credit: Miyako Nakamura/GettyFor drug discovery, protein-folding models such as AlphaFold have a problem: there aren’t enough data in public databases. Improving the performance of these artificial-intelligence-based tools will require extra data that provide examples of how proteins and drugs interact, some scientists argue.Protein structures — locked away by the thousands in drug company vaults — offer one promising source. Today, a consortium of pharmaceutical companies reports that using such data to train AI models of protein folding improves model performance markedly.The group used OpenFold3 — an open-source replication of AlphaFold 3 — to develop a new model trained on more than 20,000 proprietary protein structures. The system outperformed both comparable ones trained on public data alone and those trained on the siloed datasets of individual firms. The study, described in a blog post, has not been peer-reviewed, and the model is not publicly available.“You add all this data, and you get a pretty big bump in performance,” says Mohammed AlQuraishi, a computational biologist at Columbia University in New York City, who was part of the effort.AlphaFold is running out of data — so drug firms are building their own versionThe findings, he says, strengthen the case for generating similar publicly available datasets to supercharge protein-folding AIs. One such project, called OpenBind and supported by up to £8 million (US$10.8 million) in UK government funding, released hundreds of new protein structures last month, with thousands more in the works.An untapped veinThe Protein Data Bank (PDB), an open repository of more than 200,000 experimentally determined protein structures, was the bedrock of AlphaFold 2’s training data. It enabled the tool to predict protein structures with startling accuracy — a breakthrough recognized with the 2024 Nobel Prize in Chemistry.The model’s successors, including AlphaFold 3, added the ability to predict how proteins will interact with other molecules, including potential drugs. But the PDB has relatively few examples of experimentally determined structures interacting with drug-like molecules — maybe just 10,000, says Paul Mortenson, vice-president for computational chemistry and informatics at Astex Pharmaceuticals in Cambridge, UK.That lack of data is a problem for drug-discovery efforts. Research has suggested that the accuracy of AlphaFold 3 and other ‘co-folding’ models — which predict the structure of proteins interacting with each other — drops off a cliff when the models are challenged to predict interactions between molecules highly dissimilar to those on which they were trained1.To make such tools more useful in drug discovery, researchers say, they need access to more data — which is why they have turned to vaults of molecular structures from pharmaceutical companies.What’s next for AlphaFold and the AI protein-folding revolutionThese data are generated during drug-discovery programmes, using techniques such as X-ray crystallography and cryo-electron microscopy. Many of the protein structures have never been deposited in public databases because they relate to proprietary drug-development efforts. The total size of these vaults is unknown, but some have estimated that they could contain more data than the PDB does.“The data that’s missing from the PDB is exactly the data that’s present in our internal data,” John Karanicolas, head of computational drug discovery at the pharma company AbbVie in North Chicago, Illinois, toldNaturelast year.Better predictionsTo test whether their data could be useful for protein-folding models, AbbVie, Astex and several other drug companies last year formed a collaboration called the AI Structural Biology (AISB) Network.It involved ‘fine-tuning’ OpenFold3 — previously trained only with PDB data — on a further 20,167 structures capturing proteins bound to potential drugs, or ligands. The structures came from five companies and were provided to the model in such a way that proprietary data remained private.The AISB study found that the extra data enhanced predictions. When tested on 1,056 protein–ligand structures that were set aside from the training data, the AISB model predicted more than half of them to a high level of accuracy. By contrast, the publicly available version of OpenFold3 achieved the same performance on just one-third of the structures, and a competing open-source model called Boltz-2 achieved around 40%. The team plans to submit a paper describing the work to a peer-reviewed journal.The fact that the AISB model also outperformed co-folding tools that were trained only on each company’s individual data highlights the benefits of pooling information, says Karanicolas.
延伸阅读

更多相关文章

2026/9/24 8:45:44

2026家庭智能化升级:从伪智能到真懂你的方案选型指南

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/24 8:45:44

Redis命令大全

redis-cli 命令详解:从连接到高阶实战 redis-cli 是 Redis 官方提供的命令行客户端,是与 Redis 实例交互最直接、最灵活的工具。它既可以在命令行模式下直接执行单条命令后退出,也可以进入交互式 REPL 模式进行连续操作,还提供了监…

2026/9/24 8:45:44

T型、π型、L型滤波拓扑选型与截止频率计算实战指南

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/24 9:55:53

机房动环监控实战:Modbus TCP温湿度变送器选型与接入指南

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/24 9:55:53

BLI实验中传感器选择与蛋白固定模式解析

一、BLI生物层干涉技术的检测原理生物层干涉技术(Biolayer Interferometry,BLI)是一种基于光学干涉原理的实时、无标记分子相互作用分析技术。实验时,一个分子固定于传感器表面的生物层,作为Ligand;另一个分…

2026/9/24 9:55:53

基于TEC与PID的激光器精密温控平台DIY实战解析

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/24 9:55:53

FX5U与汇川伺服Modbus-RTU通信实战指南

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/24 9:50:52

RK3588本地部署DeepSeek大模型:Ollama与RKLLM NPU加速实战

/* MD / 富文本中的 .toc(含博客园搬家等嵌套结构);.toc-box 在侧栏,不受影响 */#content_views .toc,/* 编辑器常在目录前后插入空 p(:empty 仍占 20px),一并去掉避免顶空隙 */#content_views.markdown_views > p:empty:has(+ .toc),#content_views.markdown_views …

2026/9/23 12:07:00

GAMP 5 基于风险的计算机化系统验证:软件分类与审计追踪实践

简介:《A Risk-Based Approach to Compliant GxP Computerized Systems》即业内熟知的GAMP 5指南,面向制药企业质量与IT合规人员、验证工程师及计算机化系统管理者,用于解决GxP法规环境下系统合规性难以科学落地的问题。文档以风险管理为主线…

2026/9/23 12:06:55

安全托管MSSP实战:从静态防御到人机协同的攻防运营与应急响应

简介:这份PPT围绕互联网业务安全托管服务展开,面向企业安全负责人、IT运维人员及关注MSSP/MSS选型的读者,重点回应传统安全过度依赖人工、碎片化静态防御难以对抗产业化攻击等痛点。资源共1个pptx文件,包体约30.63MB,以…

2026/9/24 0:00:21

基于YOLOv8的渔船作业监控系统:从环境搭建到边缘部署全流程

简介:这是一套面向计算机、人工智能、自动化等专业学生与教师的毕业设计级项目资源,围绕YOLOv8实现渔船作业监控系统,可用于毕设、课程设计、大作业或项目立项演示。压缩包共97个文件,约24.21MB,以70个Python源码文件为…

2026/9/24 0:00:21

单细胞注释实战:基于Scanpy的标记基因与参考映射流程解析

简介:一份基于单细胞RNA测序数据的细胞类型注释算法研究Python毕业设计源码,针对计算机相关专业正在做毕设或需要项目实战的学习者,可用于课程设计与期末大作业。项目代码完整、经导师指导评审通过,可直接运行,覆盖数据…

2026/9/24 0:00:21

C#源生成器实战:用增量生成器替代反射,告别AOT崩溃

第一次在项目里被反射卡住,是在一个老旧的WinForms模块里:几十个类依赖PropertyChanged通知,运行时反射读属性、发通知,每次启动慢半拍不说,一上.NET Native/AOT裁剪模式几乎全面崩盘。后来我把这段逻辑全部改成C#源生…

2026/9/22 16:34:32

USB Type-C PCB布局分区设计:电源、高速信号与PD协议全攻略

做硬件这行,Type-C接口算是典型的“看着简单,做起来全坑”的东西。光引脚就24个,高低速信号、电源、控制线全部塞在一个小小的连接器里,如果PCB布局不做规划,打样回来基本就是“插上没反应”、“高速掉线”、“静电一打…

2026/9/22 20:01:30

系统编程学习原型如何补齐稳定性边界

系统编程学习原型如何补齐稳定性边界预算有限时&#xff0c;我先优化明显多余的复制&#xff0c;而不是猜测性地换容器。用借用传递只读数据通常就能减少分配&#xff1a; fn parse(line: &str) -> Result<Item, Error> { /* ... */ }用基准确认热点确实在分配&am…

2026/9/22 13:25:41

雨花区哪家财务公司代理记账比较好?

在雨花区&#xff0c;企业处理财税事务常常面临诸多挑战&#xff0c;选择一家靠谱的财务公司至关重要。湖南巨勤财务管理咨询有限公司就是本地正规实体财税服务机构&#xff0c;深耕本地工商财税行业多年&#xff0c;熟悉当地工商局、税务局最新政策与申报流程。主营公司注册、…

还想了解更多?直接咨询顾问

免费诊断 + 免费方案 + 透明报价。

全国咨询热线400-8866-253
免费获取方案
咨询二维码