New Funding - New Opportunities

OpenADMET is tackling the persistent data bottlenecks in predictive drug discovery. With major new backing, we are generating prospective ground-truth and structural data at scale. Together, we are turning ADMET from an unpredictable gamble into a rational design problem.

Share
New Funding - New Opportunities

Author: Pat Walters
DOI: pending

When we launched the OpenADMET initiative, our objective was straightforward yet fundamentally challenging: eliminate the persistent data bottlenecks that have hindered predictive absorption, distribution, metabolism, excretion, and toxicity (ADMET) modeling for decades.

Promising drug candidates fail for two main reasons. The first is lack of efficacy, which usually comes from an incomplete understanding of the underlying biology. The second is poor pharmacokinetics, unexpected toxicity, active efflux, or metabolic instability. Efficacy problems are specific to each target, so each team tackles them for the diseases it works on. ADMET problems are different: they're a public-goods problem. Any improvement in predicting them benefits drug discovery as a whole, not just a single therapeutic area or company. Yet the industry has long relied on fragmented, siloed data, retrospective benchmarks prone to data leakage, and trial-and-error medicinal chemistry heuristics.

Through a convergence of new and expanded backing from ARPA-H, the Gates Foundation, the OpenAI Foundation, and Radial/Astera Institute, OpenADMET is scaling across every dimension of ADMET. Together, these initiatives enable us to generate prospective, high-resolution physical ground truth at unprecedented scale while building the open-source infrastructure needed to separate genuine predictive signal from hype.

Where we are today

It's worth pausing to see how far we've come. In our first phase, we reduced the cost of in vitro ADMET assays by up to 100-fold, measured more than 150,000 compound–target pairs, and solved more than 400 new crystal structures of compound-bound proteins. We did all of this openly. Our datasets have been downloaded more than 30,000 times, and a community of more than 800 global researchers has grown around the work.

Here are just a few examples of how we’re improving our understanding and predictive capabilities of ADMET from the past ~2 years:

Bringing the Community Together through Blind Challenges

I am especially proud of the blind challenges. We've now run three of them, with roughly 500 participants in total. Each one asked the community to predict measurements nobody had seen yet, and then scored those predictions against new experimental data. This is the discipline that CASP brought to protein structure prediction, ultimately leading to AlphaFold. Whereas retrospective benchmarks tend to reward memorization, prospective challenges reward models that can actually predict something new. Each challenge has taught us something, both about which methods hold up and about how to design a better benchmark. Our fourth challenge is underway now, and we'll keep running challenges regularly from here on.

Connecting function to structure: PXR

The pregnane X receptor (PXR) shows what becomes possible when assay data and structures are generated side by side at scale. PXR is a nuclear receptor that senses foreign compounds and switches on CYP3A4 and other drug-metabolizing enzymes. Although its binding pocket is large and promiscuous, our structural studies have identified consistent interaction patterns. We've now generated more than 11,000 PXR assay values and 184 X-ray structures.

With data at that density, we can begin answering questions that were previously out of reach. When a small structural change turns PXR activation on or off, we can often see why in the crystal structure: a shift in ligand binding, a side-chain movement, or a change in part of the pocket. We're starting to explain structure-activity relationships instead of just cataloging them. Ultimately, we want to move beyond the black-box models we use today to clear insights that teams can use to mitigate liabilities. 

Doing the same for the CYPs

We're taking the same approach with the cytochrome P450 enzymes, which are responsible for most small-molecule drug metabolism. We have developed low-cost assays that enable us to profile thousands of diverse molecules, rather than the few hundred typically found in literature datasets. Combining those measurements with structural data allows us to see how specific interactions in the active site drive inhibition and metabolic turnover. It also lets us test our models against that ground truth.  High-throughput chemistry experiments enable us to go deeper into SAR and are again proving valuable for probing CYP structure-activity relationships (SAR), as they did with PXR. 

What comes next

The new funding lets us extend this approach across all of ADMET:

  • Metabolism (ARPA-H, Phase II). We're expanding from 10 to 25 key assay targets. We will screen these against thousands of compounds and expand structural coverage of metabolic enzymes and nuclear receptors.
  • Toxicity (Gates Foundation). We'll generate large, standardized functional datasets and structures for the hERG channel and the aminergic GPCRs. There are currently only 3 ligand-bound hERG cryo-EM structures in the PDB. We plan to expand this by 10-100 times.  Beyond hERG, we are also exploring the structures and activity of aminergic GPCRs.  While these important targets appear on every secondary pharmacology panel, there is very little standardized public data on them
  • Transporters and the blood–brain barrier (OpenAI Foundation). We'll profile tens of thousands of molecules for transport and efflux, including P-glycoprotein, and significantly increase the number of high-resolution transporter structures.
  • Blind challenge infrastructure (Radial and the Astera Institute). We're turning the tools and infrastructure we built for our own blind challenges into open, reusable infrastructure that can be used by anyone

Making ADMET a Design Problem, not a Gamble

In our recent preprint, “Mapping the Avoid-ome”, we argued that the proteins responsible for most ADMET liabilities form a finite set. Perhaps 50–100 anti-targets account for most of the problems. That means they can be mapped systematically, like any other defined part of biology. The funding described here lets us do that mapping at scale.

The bigger payoff comes from connecting the pieces. A compound's PXR activation, CYP inhibition profile, hERG liability, and P-gp efflux are not independent; they depend on many of the same molecular properties and recognition patterns. When we measure all of these across shared, overlapping compound sets, pair the measurements with structures, and test our models prospectively, we can build models that learn across the whole Avoid-ome instead of one endpoint at a time. The goal is multi-parameter design, where a chemist can anticipate a liability, understand its structural cause, and design around it before a molecule is ever synthesized.

PXR shows that this approach works. Pairing dense assay data with structures turns a hard, promiscuous target into one we can reason about. Now we need to do the same for every target in the Avoid-ome. The momentum over the past year has been remarkable, but we're still at the beginning. Looking ahead, here's what I see:

A complete, open map of the Avoid-ome. Over the next decade, I'd like every major anti-target to have what PXR has now. That means the CYPs, the nuclear receptors, the key transporters, hERG and the other cardiac ion channels, and the off-target GPCRs and kinases that appear on safety panels. Each would have tens of thousands of consistent measurements and hundreds of structures, all of which are public. The Protein Data Bank has been a public resource for more than 50 years, and it's what made AlphaFold possible. ADMET needs the same kind of shared, trusted foundation that anyone can build on.

Models that reason from protein structure, not just chemical similarity. Most ADMET models today are pattern matchers. They work well close to their training data and fail quietly outside it. Dense assay data paired with structures lets us train models that learn why a molecule hits an anti-target, not just which known molecules it resembles. That's what will let models generalize to new chemical matter, which is where drug discovery programs spend their time.

From anti-targets to patients. Protein-level measurements are the building blocks. The harder and more important problem is connecting them to what happens in people: clearance, exposure, drug–drug interactions, and toxicity. As the Avoid-ome map fills in, we need to build the bridge from the mechanism to in vivo and clinical outcomes. Then a prediction about CYP3A4 or P-gp becomes a prediction about dose and safety. Reliable, mechanistic in vitro and in silico methods are also what regulators need to move toward new approach methodologies and reduce reliance on animal testing.

Closing the loop. The pieces we're building now are cheap assays, fast crystallography, active learning, and prospective benchmarks. Together they can form a continuous cycle: models choose the most informative next experiments, the lab runs them, and the models improve. In the long run, generative design tools should treat the Avoid-ome as a set of constraints from the very first step, so molecules are designed to avoid liabilities rather than screened for them afterward.

Leveling the playing field. Large pharmaceutical companies have decades of internal ADMET data. Academic labs, startups, and groups working on neglected and global-health diseases usually don't. Open data and open models give those teams the same starting point. Some of the most important medicines of the next twenty years may come from groups that could never have built this infrastructure on their own. 

Blind challenges as permanent infrastructure. CASP has run for more than 30 years and has become the standard by which the structure-prediction field measures progress. We want OpenADMET challenges to play the same role for ADMET: a regular, trusted, independent test of which methods actually work, so claims of progress can be checked rather than taken on faith.

If we get this right, success will be easy to see. Fewer molecules will fail in the clinic for reasons we could have predicted, and ADMET will become a design problem instead of a gamble.

Join us

We can't do this alone, and we shouldn't try. Every dataset, structure, and model we produce will be released openly. We'll keep running blind challenges so the whole field can see which methods actually work. We now have the momentum and the resources to carry this work into next year and beyond. If you build ADMET models, run assays, or just want better tools for designing medicines, we'd love for you to join us.

Acknowledgements

We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute (https://ror.org/00ydx1s47)), Schrödinger Inc., the Gates Foundation, and the OpenAI Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.

This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.