Accelerating into Phase II: Mapping Drug Metabolism with ARPA-H
When we launched the OpenADMET initiative, we set out with an ambitious roadmap: to slash assay costs, generate large-scale prospective datasets, and establish a repeatable framework for evaluating ADMET machine learning models.
Thanks to an extraordinary effort from the consortium team and our collaborators, we are moving into Phase II of the AVOID-OME.
What We Accomplished in Phase I
Over the past year, the team focused on eliminating data bottlenecks that have held back predictive ADMET for decades. A few highlights from Phase I:
- High-Throughput Assays at Scale: We dropped in vitro assay costs by up to 100-fold and measured more than 150,000 compound-target pairings.
- Structural Biology: Solved over 400 new compound-bound structures, generating rich 3D structural data for critical targets.
- Variant Profiling: Validated a VAMP-seq approach to assess protein stability across every cataloged variant of key metabolic enzymes and receptors.
- Open Tools and Community Engagement: Released our open-source ML infrastructure and engaged roughly 500 blind challenge participants, over 30,000 dataset downloads, and an active community of 800+ researchers on Discord.
Sharpening the Focus for Phase II: Tackling Metabolism Head-On
In Phase II, we are directing our resources into a deep, comprehensive map of drug metabolism. Metabolic interactions account for roughly 30% of adverse drug reactions and 10 to 15% of clinical failures. Unlike some broader phenotypic ADMET endpoints, drug metabolism is largely governed by a well-defined, tractable set of enzymes and nuclear receptors. It is one of the rare instances in drug discovery where a massive problem is also scientifically addressable with the right data.
In Phase II, OpenADMET will:
- Build and screen 15 key assay targets against tens of thousands of diverse small molecules.
- Expand structural coverage across metabolic enzymes and nuclear receptors.
- Continue quarterly blind challenges and open data releases to benchmark predictive models under genuine prospective conditions.
The data coming out of our pipelines is already challenging standard medicinal chemistry assumptions about CYP liabilities, showing how much mechanistic structure-activity relationship (SAR) remains to be uncovered. Phase II will give the entire field the data and structure needed to separate real predictive signal from rules of thumb that fail in the clinic.
Moving Fast with ARPA-H
We are deeply grateful to ARPA-H for their support and operational agility. Their funding model moves quickly, stays disciplined on real-world impact, and allows teams the flexibility to pivot where the science leads.
To everyone participating in our blind challenges, downloading datasets, and contributing to the open discussions: thank you. We have a lot of exciting data and challenges lined up for the coming months.
You can follow our updates, download datasets, and join upcoming challenges at openadmet.org.
We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute (https://ror.org/00ydx1s47)), Schrödinger Inc., the Gates Foundation and the OpenAI Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.
This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.