Announcing OpenADMET’s CYP inhibition blind challenge
Who will be the Heme-coming queen?
Introduction
OpenADMET facilitates blind challenges to bridge the gap between computational modeling and experimental reality in drug discovery. By providing high-quality, blinded datasets, these challenges offer a transparent benchmark for evaluating the predictive power of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) models. Building on the success of our PXR induction challenge, and with sincere thanks to everyone who participated, we are moving one level down the biological cascade. PXR and other nuclear receptors are the xenobiotic sensors that marshal the drug-metabolizing machinery; the cytochrome P450s are that machinery
Cytochrome P450 (CYP) enzymes are the primary drivers of xenobiotic metabolism. This heme-containing superfamily is responsible for the biotransformation of the vast majority of marketed small-molecule drugs and contains dozens of human isoforms with overlapping and complementary substrate scopes. This flexibility is exactly what makes them so central to xenobiotic metabolism and makes their collective behavior difficult to predict. There is a vast array of clinically relevant CYPs, but for this challenge, we have chosen to focus on four isoforms that are the core drivers of small-molecule drug metabolism: CYP3A4, CYP2C9, CYP2D6, and CYP1A2. Other CYPs (e.g., CYP2C19, CYP2C8, and CYP2B6) will come later in our journey.
This challenge is once again made possible by our in-house data generation efforts at Octant and our structural biology capabilities at UCSF, in collaboration with the Fraser lab. We’re beginning to hit our stride with data generation, and expect data to drop at an accelerating pace from here! We are passionate about getting our data into the community’s hands as fast as possible so that we can make collective progress. So start your engines, submissions open in ~2 weeks!
Why CYP Inhibition Matters
Where PXR induction accelerates metabolism, CYP inhibition does the opposite, jamming the body’s mechanism for xenobiotic elimination. Both have the potential to derail drug discovery programs through the same general mechanism: unpredictable drug exposure.
When a candidate compound inhibits a CYP isoform, it slows the clearance of any co-administered drug that depends on that enzyme, driving plasma concentrations upward, sometimes into toxic territory. The consequences include:
-
Drug–Drug Interactions (DDIs): Reduced clearance of co-administered drugs can push exposure above the therapeutic window, the most common and clinically significant CYP-mediated liability.
-
Narrow therapeutic index risk: Even modest inhibition of a key isoform can be dangerous when the affected co-medication has little margin between efficacy and toxicity.
-
Time-dependent inhibition (TDI): Some drugs that are CYP substrates yield metabolites that show increased inhibition compared with the drug itself. In TDI, the drug is converted into either a more potent reversible inhibitor, or a mechanism-based inhibitor, which is a reactive intermediate that forms covalent bonds with either the protein or heme cofactor. The latter irreversibly inactivates the enzyme, producing DDIs that persist beyond the compound's own presence.
Consequently, evaluating CYP inhibition liabilities is a fundamental pillar of the ADMET cascade and is a core expectation of the FDA and other regulatory agencies. The FDA publishes a detailed table that covers marker inhibitors and substrates for these interactions, as well as regulatory guidance on evaluating and measuring these risks in the clinic.
Predicting inhibition across this panel of isoforms in silico would enable teams to prioritize cleaner candidates, design out interaction liabilities before they become costly late-stage failures, and reduce the need for clinical DDI studies in late-stage development. Approaches like these align well with the FDA’s New Approach Methodologies (NAMs) guidance, which aims to reduce animal testing in preclinical development.
While inhibition data exists in the public domain, high-quality, consistently generated data across a defined isoform panel remains rare, and this is the gap this challenge is built to fill. Like PXR, the CYPs present a modeling challenge that goes beyond a single number. Their active sites are large, malleable, and heme-centered, accommodating chemically diverse substrates and often binding in multiple orientations (and multiple times). As with PXR, structural biology provides a valuable lens here, helping us rationalise observed SAR across our datasets.
Developing the assay
As with our PXR challenge, we developed a high-throughput, in-house CYP assay panel compatible with Octant’s Navigator platform.
We will share a more thorough description of our assay development efforts in a separate blog post within the next few weeks. We first developed CYP inhibition assays adapted from commercially available fluorescence-based Vivid CYP inhibition assay kits from ThermoFisher. Our inhibition assay involves the use of masked probes, which can be metabolized by CYPs to produce a fluorescent product. Using different probes for each CYP isoform (CYP3A4: DBOMF, CYP1A2 and CYP2C9: EOMCC ), we scaled CYP3A4, 2C9, and 1A2 inhibition assays down to a format compatible with 1536-well plates. These plates are read out via fluorescence either at a single concentration as a primary screen, or as a dose-response curve (DRC) to generate IC50 (half maximal inhibitory concentration) estimates. Validation of these assays showed acceptable signal-to-noise ratios and dynamic range. Note that unlike the assays in our PXR challenge, these experiments are biochemical rather than cell-based.
CYP2D6 was the one major isoform where the fluorogenic approach to measure inhibition was ineffective. The available fluorescent 2D6 probes gave poor signal-to-noise ratios and dynamic ranges in our hands. Rather than ship a low-confidence 2D6 assay, we pivoted to a label-free mass-spec readout that directly measures substrate turnover. We replaced the masked fluorogenic probe with a drug-like 2D6 substrate (dextromethorphan) and quantified parent depletion on a SCIEX Echo-MS system, which we have already been using at OpenADMET for high-throughput measurements of CYP metabolism. Reactions are run with the same recombinant enzyme and NADPH-driven biochemical format as the fluorescence panel, with parent depletion monitored to determine inhibition and estimate IC50 values. Briefly, the Echo-MS uses acoustic ejection mass spectrometry (AEMS): acoustic droplet ejection fires nanoliter-scale droplets (~2.5 nL) from each well into an open-port interface that dilutes out matrix effects and carries the sample to a SCIEX ZenoTOF 7600 mass spectrometer for analysis, with no chromatography step needed. Contactless sampling with low carryover and good tolerance to suppressive buffers gave us the throughput and data quality we needed for both single-concentration primary screening and the DRC follow-up, while also using a standard DDI probe for CYP2D6, dextromethorphan, rather than a surrogate fluorophore.
Beyond reversible inhibition, we also profile each compound for time-dependent inhibition (TDI), the hallmark of mechanism-based ("suicide") inhibitors, whose potency grows as the CYP metabolizes them into reactive species that reversibly, quasi-irreversibly, or covalently disable the enzyme. Because TDI drives some of the most serious and hardest-to-predict clinical drug–drug interactions, our DRC protocol collects a TDI curve for every compound by default.
We assess TDI using the standard IC50-shift design (Figure 1). Each compound is pre-incubated with the recombinant CYP in the presence (TDI condition) and absence (direct inhibition condition) of NADPH before the probe substrate is added and residual activity is read out. A compound that is only a reversible inhibitor gives the same IC50 in both arms; a time-dependent inhibitor is converted by the enzyme during the NADPH pre-incubation into a more potent (often irreversible) inhibitor, so its dose-response curve shifts left, to a lower IC50, in the +NADPH arm. The magnitude of that leftward shift is the TDI signal. For both the fluorogenic and MS assay, we benchmarked both the reversible inhibition assay and the TDI assay against known reference inhibitors, which we will detail in the forthcoming data-focused blog post. In summary, the benchmarking experiments confirmed each assay's accuracy and fidelity.

Figure 1: IC50-shift design allowing assessment of both direct inhibition and TDI. Compound is pre-incubated with recombinant CYP ± NADPH before probe substrate addition and readout of residual activity. Compound A gives superimposable curves; Compound B shifts left in the +NADPH arm.
How we constructed the dataset
We built our CYP inhibition dataset using a two-tier screening strategy followed by hit expansion, mirroring the approach we used for our PXR challenge dataset.
Starting from two compound libraries, Enamine DDS10 (a diversity set) and FDAA (FDA-approved compounds), we ran a single-concentration primary screen against each CYP. Note that the single concentration screen was run in the active preincubation (TDI) condition to catch as many positives as possible. To select hits for dose-response confirmation, we performed a variance-stabilizing transformation, corrected for spatial artifacts, fitted a per-plate linear model that contrasted each compound against the negative control, and applied a Benjamini-Hochberg correction to control the false discovery rate (FDR). Compounds were then selected as hits based on log-fold change and FDR thresholds. Because the assay measures inhibition, a negative log-fold change indicates that a compound inhibits the CYP.
For each hit, we collected 12-point dose-response curves and fit them using a Bayesian curve-fitting approach, targeting roughly 1,500 DRCs per CYP (see the PXR induction post for more details on the curve-fitting procedure). Similarly to the single-dose data, we stabilized the variance and corrected for spatial artifacts prior to fitting the DRCs. It’s worth noting that in our DRC protocol, TDI curves are obtained for each compound by default.
After evaluating the DRC dataset, we carried out hit expansion on the most potent compounds to create our test set, selecting the top 25 hits per CYP for 3 CYPs (75 compounds total) and purchasing the top 10 chemisimilars (ranked by Tanimoto similarity) per hit compound from the Enamine US in-stock catalog. This hit expansion was conducted over CYP1A2, CYP2C9, and CYP3A4 hits only, as the CYP2D6 assay was still in development on the Echo-MS system. This led to a total of 750 compounds, which were then assayed in DRC mode against each of the four CYPs (including CYP2D6, see Figure 2 for a visual breakdown). The resulting DRC training set is a sparse matrix, while the test set is dense across the CYPs.

Figure 2: Schematic of the train-test split for the OpenADMET CYP inhibition competition
A more detailed breakdown of how we developed our CYP assay capabilities, including the assay itself and full dataset construction, is coming in a follow-up blog post. Cumulatively, this leaves us with a fascinating dataset to both evaluate machine learning models and examine cross-CYP inhibition and reactivity patterns (another focus of ours).
Challenge details
As usual, our competition will be run on a dedicated Hugging Face space with a Discord channel for Q&A, announcements, and general support.
We are reverting to a simpler challenge structure for our upcoming CYP challenge, compared with our previous PXR challenge. There will be only one stage, with no release of half the test set halfway through this time. While instant gratification may have been good for people's attention span (kids these days), we view the simpler challenge design as appropriately difficult with a lighter operational overhead.
As a reminder, due to the design of the assay to enable detection of TDI, each compound has two arms: the direct inhibition arm with a -NADPH preincubation and the TDI arm with a +NADPH preincubation. The direct inhibition arm is what people normally think of as direct, reversible inhibition. Because no NADPH is present during the preincubation, the enzyme cannot turn the compound over, so no metabolism-dependent inactivation can occur, and any inhibition observed reflects the parent compound binding directly to the enzyme. The TDI arm, by contrast, permits catalytic turnover during the preincubation, so it captures both this direct inhibition and any additional time-dependent inhibition arising from reactive metabolites or metabolic intermediates that covalently (or quasi-irreversibly) inactivate the enzyme.
The challenge is split into two tracks that reflect how these assays are actually used in drug discovery.
Direct inhibition (regression). Participants predict the direct-inhibition pIC50 for each of the four CYP isoforms CYP3A4, CYP2D6, CYP2C9, and CYP1A2 for the 750-compound test set, for a total of 4 regression targets per compound. Low activity compounds will be downweighted in the evaluation function (see later).
Time-dependent inhibition (classification). Rather than predicting the TDI-arm pIC50 directly, participants classify whether a compound is a time-dependent inhibitor, i.e., whether the IC50 shift after preincubation exceeds 2-fold (a boolean True/False label). This mirrors how the assay functions in practice: as a prescreen whose positive calls trigger a detailed follow-up kinetic study, with a ~2-fold shift being a common threshold that would trigger a follow-up. Predictions are requested for every compound in the test set, so that the set of scored compounds carries no information back into the regression task, but only compounds whose label can be assigned with confidence contribute to the score.
Compounds with a direct-inhibition pIC50 above 4 are labeled positive or negative according to whether the TDI shift exceeds 2-fold. Compounds falling below a pIC50 of 4 in the direct-inhibition arm are labeled positive when the TDI-arm pIC50 exceeds 4.301, since a true direct-inhibition value of at most 4 guarantees a greater than 2-fold shift in that case (log10(2-fold) = 0.301), known here as an inferred positive. These compounds correspond to going from no signal in direct inhibition to measurable TDI and we believe are important to include. When the pIC50 falls below 4 for both arms we cannot reliably determine TDI, but at such low levels of activity, these compounds are unlikely to be selected for a follow up assay, and are known here as assigned negatives. The positive class is positives + inferred positives (red) and the negative class is negatives + assigned negatives (green). See Figure 3 for a visual breakdown.

Figure 3: Classification scheme for the TDI track.
Because companies overwhelmingly screen for TDI against CYP3A4 and only rarely against other isoforms, and because we observe very few shifts for CYP1A2 or CYP2C9, TDI classification is evaluated only for CYP3A4 and CYP2D6.
We will provide an extensive training data pack including primary screening data for the whole chemical library (in the TDI condition) and ~1,500 DRCs per CYP isoform, covering both arms so participants can derive shift labels for training.
Evaluation
For the direct inhibition track, the primary metric will be a modified macro-averaged relative absolute error function with respect to the test set, averaged across the endpoints (MA-ST-RAE, see later), with a battery of other metrics available, assessed across all four isoforms. Low activity compounds will be downweighted.
For the TDI classification track, submissions are scored by Matthews Correlation Coefficient (MCC), which is well suited to the imbalanced, threshold-based nature of the shift labels.
Half of the test set will be used for a live leaderboard, split by chemisimilar series, such that all compounds from a parent end up in either the live leaderboard or the fully blinded set. There will be an interim leaderboard at the halfway mark, at which participants' performance on the full test set will be revealed only once.
Further detailed information will be provided on launch day
Timeline, evaluation, and changes to competition rules
Running these challenges as an organising team has been a constant learning exercise. In addition to participants pushing the field forward, we as an organizing team also aim to continue to innovate in the preparation, execution, and evaluation of our challenges. Taking on board feedback from our previous challenges, as well as operational constraints we have encountered, we are making changes to our evaluation, competition operations, and rules.
-
We will have an award for the most innovative machine learning approach(es). In addition to our standard leaderboard winners, we will recognize the most innovative machine learning approach(es), as judged by the OpenADMET team. At OpenADMET, we want to create a platform for innovation in computational methods, alongside great datasets. For this to work, we need to think carefully about our objective function when running our blind challenges. As such, this award exists to reward exploration rather than only exploitation of what we as a community already know works. We will weigh factors such as architectural novelty, creative use of available data, novel training or uncertainty quantification strategies, simple yet effective architectures, and ideas that are interesting even if they don't top the leaderboard. Importantly, eligibility for this award is somewhat decoupled from leaderboard rank; a lower-scoring entry can still win on the strength of its ideas, so please don't let the fear of a mediocre score stop you from trying something ambitious. Final judgment is at our discretion, winners will be invited to speak at one of our webinars!
-
We will downweight low-activity compounds. In our previous PXR challenge, many participants noted that the low-activity compounds below (pXC50 < 4) were the “widow makers” and proved very challenging for participants to predict with any certainty. These low-activity values are below the lowest tested dose in our assay and, as such, are difficult to determine with precision, meaning measuring performance on these compounds is not very informative for model evaluation. To remedy this, we are using a custom Soft-Threshold Relative Absolute Error (ST-RAE) scoring function that explicitly accounts for ground-truth uncertainty. In addition to point estimates, our fitted DRCs yield credible intervals for each pIC50. Empirically, these intervals expand significantly at lower activity levels, aligning with assay precision limits. Under this metric, error is measured as the distance between a predicted value and the nearest bound of the credible interval; predictions falling anywhere inside the credible interval incur zero error. This soft-threshold approach ensures our leaderboard scores accurately reflect true model performance without penalizing predictions for falling within experimental noise.
-
We are no longer committing to writing a summary preprint for each challenge. We believe open science is best served by high velocity: getting well-curated data and clear results into the community's hands quickly, and often, rather than gating them behind the long tail of writing full preprints. Given finite time and resources, we'd rather run more challenges and release more data than slow the cadence to produce a preprint for each one. We will be writing up our results as blog posts and diving deep into performance, but in shorter form. Participants remain free and encouraged to publish their own methods and findings. Our previous commitments to preprints from the ExpansionRx and PXR challenge still stand.
-
We will not accept multiple leaderboard submissions from the same team/lab. Previously, in our PXR challenge, we allowed multiple submissions from within the same lab/team where the “approaches differed significantly”. Understandably, people’s definition of “differed significantly” is subjective. To avoid confusion, we are eliminating the ability to submit separately more than once, but as always, rely on honesty from our participants. Team/Lab refers to “a collection of people who cooperate intensively to prepare a submission” in this context. Judgment is at our discretion.
-
As with our PXR challenge, we require disclosure of the use of proprietary data. While large volumes of PXR agonism data are perhaps rarer in the vaults of large companies, we hypothesize that CYP inhibition data is not, and will present a significant advantage to those that possess it (see results of the ExpansionRx challenge). As such, we require participants to disclose any use of proprietary data so that we can compare contestants fairly.
-
We will also track open code. We will have a checkbox for participants to record if they have code available on the open internet to reproduce their results (there can be a slight delay, but a GitHub repo with some code is expected). This will win you kudos from us and from the community and is especially important if you are wanting people to notice your work on novel machine learning ideas.
The challenge timeline is shown below; times are one minute to midnight UTC unless otherwise specified:
- Challenge announced: 2026-07-29
- Challenge launch: 2026-08-17
- Submission deadline for intermediate leaderboard: 2026-09-24
- Intermediate leaderboard release: 2026-09-25
- Challenge close, deadline for final submissions: 2026-11-03
- Webinars, blog, and wrap up after 2026-11-03
We are very excited to bring you our next challenge and look forward to seeing what you can do with the data!
The OpenADMET Team
Acknowledgements
We would like to acknowledge the hard work of all the experimentalists at Octant and UCSF for making this challenge possible. In particular, we would like to thank Lauren Orr, Ayesha Ghazali, Scott Simpkins, Ana Lindahl, Terry Lou, Sam Sabaat, Marco Falcioni, Bryan Jiang, and Robert Warneford-Thomson from Octant as well as Karen Orta, Galen Correy, and others from the Fraser lab at UCSF.
We would like to thank our funders for their support of OpenADMET, in particular ARPAH, Radial (part of the Astera Institute (https://ror.org/00ydx1s47)), Schrödinger Inc, and the Gates Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.
This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.