Halfway Through the OpenADMET CYP Challenge
The Heme-coming last dance begins: Interim leaderboard release & announcing a new Structure Track
Authors: Maria Castellanos, Jon Swain, and Hugo MacDermott-Opeskin
DOI: Pending
Can you believe we are already halfway through the challenge?! The level of community participation so far has been mind-blowing. The OpenADMET challenge community never fails to amaze us!
Today we are releasing the intermediate leaderboard, a one-off chance to see how your models perform across the entire test set. As a reminder, the leaderboard you’ve seen until now is a live leaderboard, evaluated on only half of the test set. The version we’re releasing today has been evaluated on the entire test set. As a result, you may notice the ranking may have shifted a bit. We do this to prevent overfitting to the public leaderboard and to highlight models that truly generalize.
In other news, today we are also launching the Structure Track, featuring 20 brand-new, never-before-seen protein-ligand structures of CYP3A4, resolved via cryo-EM by our team at UCSF!
The leaderboards
Here’s where things stand at the halfway mark. Below are the first 19 entries for the Direct Inhibition leaderboard, and the first 20 entries of the TDI leaderboard. You can explore the full interim leaderboard on our Hugging Face space.
A few things to note:
We removed entries suspected of violating challenge rules.
We have been a little disappointed by a small number of participants who have been trying to circumvent our challenge rules. As a reminder, only one entry per team is allowed. Our system automatically overwrites multiple entries with the same username, keeping only the latest one.
Using alternative accounts to circumvent this rule is not allowed, and we reserve the right to remove entries and block users that we suspect are engaging in these actions. If you believe your entry was omitted by mistake, reach out to us via email.
We use a tier-based system to group entries that are not significantly different from each other.
To account for statistical noise, we group submissions that perform similarly into tiers using the Benjamini-Hochberg procedure (detailed in our previous blog post). Starting with the top entry, subsequent submissions are grouped into Tier 1 until we encounter a model whose performance differs significantly, which is then assigned to Tier 2. In the Direct Inhibition track, only one entry made it into Tier 1, with the second entry being significantly different (Tier 2). For TDI, a whopping 26 entries were assigned to Tier 1 and are therefore indistinguishable according to this test. Below, we show the top 18 entries from the Direct Inhibition leaderboard (spanning Tiers 1–4) and all 26 Tier 1 entries from the TDI track..
The interim leaderboard might be a little different than the live leaderboard.
Because the live leaderboard reflects only half the test set, your rank may have moved up or down. The interim leaderboard provides a much clearer picture of what to expect in the final standings, so use this benchmark to fine-tune your approach for the home stretch!
Direct Inhibition Interim Leaderboard
| Rank | Username | Significance | Model report | Proprietary data? | Open-source code? | MA-ST-RAE | MA-MAE |
|---|---|---|---|---|---|---|---|
| 1 | preheat-to-450 | Tier 1 | NA | ✘ | ✘ | 0.3721 | 0.6005 |
| 2 | wbot | Tier 2 | NA | ✘ | ✘ | 0.3912 | 0.6195 |
| 3 | cypadeedoodah | Tier 2 | NA | ✘ | ✘ | 0.3921 | 0.6511 |
| 4 | Rubyy | Tier 2 | NA | ✘ | ✘ | 0.4014 | 0.6394 |
| 5 | EleanorRigby | Tier 3 | NA | ✘ | ✘ | 0.4173 | 0.6627 |
| 6 | beast | Tier 3 | NA | ✘ | ✘ | 0.4263 | 0.6947 |
| 7 | leebrubeck | Tier 4 | NA | ✘ | ✘ | 0.4336 | 0.6264 |
| 8 | hfger | Tier 4 | NA | ✘ | ✘ | 0.4357 | 0.6383 |
| 9 | first-chair | Tier 4 | NA | ✘ | ✘ | 0.4362 | 0.6306 |
| 10 | furtivepeony | Tier 4 | NA | ✘ | ✘ | 0.4370 | 0.6383 |
| 11 | Scigantic | Tier 4 | NA | ✘ | ✘ | 0.4382 | 0.6568 |
| 12 | psk | Tier 4 | NA | ✔ | ✘ | 0.4387 | 0.6424 |
| 13 | hf-far | Tier 4 | NA | ✘ | ✘ | 0.4410 | 0.6680 |
| 14 | rasayan-labs | Tier 4 | link | ✘ | ✘ | 0.4413 | 0.6449 |
| 15 | briford | Tier 4 | link | ✘ | ✔ | 0.4430 | 0.6415 |
| 16 | horse-pusher | Tier 4 | NA | ✘ | ✘ | 0.4440 | 0.6689 |
| 17 | as-ris-1 | Tier 4 | NA | ✘ | ✘ | 0.4444 | 0.6425 |
| 18 | curiousone | Tier 4 | NA | ✘ | ✘ | 0.4469 | 0.6447 |
TDI Interim Leaderboard
| Rank | Username | Significance | Model report | Proprietary data? | Open-source code? | MA-MCC | MA-Accuracy |
|---|---|---|---|---|---|---|---|
| 1 | nova | Tier 1 | NA | ✘ | ✘ | 0.3980 | 0.8552 |
| 2 | CYPblend | Tier 1 | NA | ✘ | ✘ | 0.3888 | 0.8387 |
| 3 | jeremy | Tier 1 | link | ✘ | ✔ | 0.3868 | 0.8451 |
| 4 | TeamPozeSCAF | Tier 1 | NA | ✘ | ✘ | 0.3847 | 0.8544 |
| 5 | N283T | Tier 1 | NA | ✘ | ✘ | 0.3798 | 0.8478 |
| 6 | stir_bar | Tier 1 | link | ✘ | ✘ | 0.3796 | 0.8397 |
| 7 | harinu123 | Tier 1 | NA | ✘ | ✘ | 0.3767 | 0.8464 |
| 8 | ADI | Tier 1 | NA | ✘ | ✘ | 0.3732 | 0.8505 |
| 9 | Asidsal11 | Tier 1 | NA | ✘ | ✘ | 0.3729 | 0.8591 |
| 10 | TeamPrescience | Tier 1 | NA | ✘ | ✘ | 0.3714 | 0.8502 |
| 11 | preheat-to-450 | Tier 1 | NA | ✘ | ✘ | 0.3653 | 0.8357 |
| 12 | lex | Tier 1 | NA | ✘ | ✘ | 0.3622 | 0.8409 |
| 13 | PIPL-Bio | Tier 1 | NA | ✘ | ✘ | 0.3609 | 0.8389 |
| 14 | EleanorRigby | Tier 1 | NA | ✘ | ✘ | 0.3607 | 0.8362 |
| 15 | newcyp | Tier 1 | NA | ✘ | ✘ | 0.3558 | 0.8162 |
| 16 | NIPERKOLKATA_PI_LAB | Tier 1 | NA | ✘ | ✘ | 0.3553 | 0.8364 |
| 17 | NorthCoil | Tier 1 | NA | ✘ | ✘ | 0.3553 | 0.8113 |
| 18 | C_CYPher-graphattn-B2regaux | Tier 1 | NA | ✘ | ✘ | 0.3543 | 0.8065 |
| 19 | curiousone | Tier 1 | NA | ✘ | ✘ | 0.3533 | 0.8349 |
| 20 | horse-pusher | Tier 1 | NA | ✘ | ✘ | 0.3522 | 0.8337 |
| 21 | yourchoice | Tier 1 | NA | ✘ | ✘ | 0.3468 | 0.8246 |
| 22 | 450nm | Tier 1 | link | ✘ | ✘ | 0.3457 | 0.8376 |
| 23 | Tutu | Tier 1 | NA | ✘ | ✘ | 0.3451 | 0.8173 |
| 24 | furtivepeony | Tier 1 | NA | ✘ | ✘ | 0.3443 | 0.8280 |
| 25 | srijitseal | Tier 1 | NA | ✘ | ✘ | 0.3437 | 0.7773 |
| 26 | aries | Tier 1 | NA | ✘ | ✘ | 0.3406 | 0.8400 |
The community has risen up to the challenge
The response from the community has been phenomenal! We’ve seen steady growth since the start of the challenge, with over 200 unique participants in the Direct Inhibition (Regression) track and over 100 in the TDI (Classification) track.
But the challenge is far from over. If you have been waiting to participate, now is the time to jump in! You still have over a month to submit your predictions before the challenge closes on November 3, 2026.
Time to work on your CYP2D6!
Here’s the trend we have seen so far: CYP3A4 has been the easiest isoform to predict across both tracks, while CYP2D6 has proven much trickier (see figure below).

To find why, we need to look at how the dataset was constructed. As detailed in our announcement blog post, the test set compounds were selected from a hit expansion on the most potent compounds on the training set. For three of the isoforms (namely, CYP3A4, CYP1A2, and CYP2C9), we took the top 25 hits and expanded each into 10 chemo-similar analogues (750 compounds total). Because CYP2D6 hits were not explicitly prioritized in this expansion, the test set for CYP2D6 contains inherently less potent compounds than the other three isoforms. This distribution shift between training and test sets makes CYP2D6 a much harder target to predict in this competition.
We’re only halfway through the challenge, so we can’t wait to see what creative strategies you come up with to boost performance on CYP2D6 over the coming weeks!
And now (*drumroll*) we introduce our new track
Today we also mark the start of a new Structure Track which will run alongside the inhibition tracks through the end of the competition.
Our team at UCSF has resolved 20 novel Cryo-EM structures of the CYP3A4 isoform complexed with a selection of assayed ligands (15 from the training set, 2 from the test set, and 3 additional compounds). Participants are tasked with predicting the 3D structure of a protein-ligand complex given only the SMILES strings.
CYP3A4 features a remarkably flexible binding pocket that undergoes significant conformational changes to accommodate diverse chemical matter. This elasticity, combined with the fact that similar ligands can adopt drastically distinct binding poses, makes CYP3A4 a nightmare for computational structure prediction. These 20 high-quality Cryo-EM structures significantly expand the chemical space beyond existing X-ray PDB structures. For more details on the experimental design, check out our recent preprint.
How to participate
You will submit predicted 3D structures for all 20 complexes compiled into a single .zip archive. Feel free to leverage co-folding models, docking software, public PDB structures, or proprietary data. For inspiration, we recommend looking into some of the workflows that were successful in our previous PXR challenge, but we are also looking forward to seeing all the creative approaches y’all may come up with!
Full instructions for loading data and formatting submissions are available in our tutorial.
Rules
Submissions will be scored primarily using the Local Difference Test for Protein Ligand Interactions (LDDT-PLI), while other metrics ,such as Binding-Site Superposed, Symmetry-Corrected Pose Root Mean Square Deviation (BiSyRSMD), will be provided as secondary metrics. Comparisons are run automatically against the blinded ground-truth using the OpenStructure pipeline. Any structure that fails basic OpenStructure checks receives an LDDT-PLI score of 0 and a BiSyRMSD penalty of 20 Å.
What’s new this time? We’re integrating PoseBusters into the scoring framework. Co-folding algorithms often optimize for raw atom placement at the expense of physical plausibility (e.g., clashing atoms or strained geometries). This was demonstrated by Buttenschoen, et. al., who designed a set of tests to validate chemical and geometric consistency of a ligand pose. We believe that accuracy metrics, such as LDDT-PLI or BiSyRSMD don’t tell the whole story ,and producing realistic poses is equally important, so every structure will now pass through PoseBusters physical/chemical sanity checks. If a prediction fails these checks, it will receive an LDDT-PLI score of 0. Look out for an upcoming blog post where we run a retrospective PoseBusters analysis on our earlier PXR challenge structures!
As with the other tracks, submission feedback will be given automatically in a dedicated Discord channel, including a list of structures that don’t pass either OpenStructure or PoseBuster checks. In addition, all the scripts we use for structure evaluation are provided in the tutorial repo, so you can evaluate your models locally prior to submission.
On to the last dance
The platform setup remains unchanged for the final phase. The live leaderboard will stay active in our Hugging Face Space until the final deadline: The platform setup remains unchanged for the final phase. The live leaderboard will stay active in our Hugging Face Space until the final deadline: November 3, 2026, at 11:59 PM UTC.
A quick checklist as we approach the end of the challenge:
- Use the interim leaderboard results to adjust your workflows for your final submission. Keep in mind that performance on the live leaderboard is not a true reflection of your performance on the final standings
- Remember that we require a comprehensive model report along with your final submission. If this report is not included by the time the submission closes, you will not be considered for the final leaderboard.
- Finally, we maintain a strict policy against multiple submissions per team; any attempts to circumvent this policy may result in disqualification at our own discretion
Good luck and happy modeling!
Acknowledgements
We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute), Schrödinger Inc, and the Gates Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.
This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.