Escape from Flatland, CYP Edition
Revisiting Merck's conclusions with our data
Revisiting Merck's conclusions with our data
At OpenADMET, open science isn't just a buzzword for us. We're tearing down silos by sharing raw data, open-source models, and runnable code instead of static papers. Through blind challenges and real-time collaboration, we're building a faster, transparent future for drug discovery.
We set out to fine-tune OpenFold3 on PXR, a notoriously promiscuous target. Along the way we found that success does not come down to how much data you have, but on whether that data consistently shows the model what it needs to see.
The Heme-coming last dance begins: Interim leaderboard release & announcing a new Structure Track
OpenADMET is tackling the persistent data bottlenecks in predictive drug discovery. With major new backing, we are generating prospective ground-truth and structural data at scale. Together, we are turning ADMET from an unpredictable gamble into a rational design problem.
After the PXR blind challenge, we found that feeding a CheMeleon embedding and primary screen predictions into a tabular foundation model drove a top entry's accuracy. The combination cut our MAE from 0.50 to 0.43 and is now a first-class pipeline in openadmet-models.
Author: Pat Walters DOI: 10.5281/zenodo.23176628 When we launched the OpenADMET initiative, we set out with an ambitious roadmap: to slash assay costs, generate large-scale prospective datasets, and establish a repeatable framework for evaluating ADMET machine learning models. Thanks to an extraordinary effort from the consortium team
Author: Pat Walters DOI: 10.5281/zenodo.23176668 Small molecules remain one of the most practical tools for reaching patients in underserved regions. They can be formulated for ambient storage, manufactured at scale and low cost, and taken as an oral pill rather than an injection — unlike many vaccines and
Ninety-six semi-pure compounds span six 12-analog series, ranked by potency and yield-corrected. Five co-crystallized PXR leads map SAR into the pocket via scaffold-aligned conformers. Inspect complexes, toggle ligands, probe polar contacts, and track alignment quality alongside 2D structures.
Announcing the open-source release of our production-ready blind challenge infrastructure template, enabling anyone to spin up a molecular prediction competition in an afternoon.
Moving into the A and D of OpenADMET
Knowing where in chemical space your models will give reliable predictions is essential to getting actionable insights. But how do you figure out where your model is applicable?
Blind Challenges
Exploring the limits of ligand-based ML models in predicting PXR induction, based on key learnings from our recent blind challenge.
Octant details their robust, reproducible, and scalable assays to map the inhibition of the most pharmacologically relevant cyotochrome P450 isoforms for >10k molecules across chemical space.
We are excited to announce that OpenADMET has received new grant funding from Radial, the life sciences division of the Astera Institute.
The heme-coming ball has begun
Medicinal chemistry groups frequently rely on mental lists of structural alerts to predict CYP inhibition, yet these heuristics are rarely tested against real data. What happens when we evaluate these alerts using 7,800 experimental measurements?
We attempted to improve the CheMeleon foundation model. We were ultimately unsuccessful, but we learned valuable lessons about foundational training along the way.
Who will be the Heme-coming queen?
Models
We extend our previous active learning analysis to choose not just which compound to test, but which assay to run. Utilizing both primary screens and full dose-response measurements cuts the cost of finding actives by more than half, without costing the model any accuracy.
News
Our quarterly newsletter detailing our progress, goals and priorities.
Blind Challenges
Announcing the results of the PXR Blind Challenge — and the public release of the largest high-quality PXR induction dataset ever made available, with over 11,000 compounds.
Models
How do you train robust deep learning models when most of the high-quality data is proprietary? Discover how to enrich multitask ADMET predictions using a public proxy for proprietary learnings.
Models
Part 1 of a series on cofolding methods for ADMET targets: using structure prediction to model protein–ligand complexes for key Avoidome anti-targets like PXR and CYP3A4.