Halfway Through the OpenADMET CYP Challenge

The Heme-coming last dance begins: Interim leaderboard release & announcing a new Structure Track

Share
Halfway Through the OpenADMET CYP Challenge

Authors:  Maria Castellanos, Jon Swain, and Hugo MacDermott-Opeskin
DOI:  Pending

Can you believe we are already halfway through the challenge?! The level of community participation so far has been mind-blowing. The OpenADMET challenge community never fails to amaze us!

Today we are releasing the intermediate leaderboard, a one-off chance to see how your models perform across the entire test set. As a reminder, the leaderboard you’ve seen until now is a live leaderboard, evaluated on only half of the test set. The version we’re releasing today has been evaluated on the entire test set. As a result, you may notice the ranking may have shifted a bit. We do this to prevent overfitting to the public leaderboard and to highlight models that truly generalize.

In other news, today we are also launching the Structure Track, featuring 20 brand-new, never-before-seen protein-ligand structures of CYP3A4, resolved via cryo-EM by our team at UCSF!

The leaderboards

Here’s where things stand at the halfway mark. Below are the first 19 entries for the Direct Inhibition leaderboard, and the first 20 entries of the TDI leaderboard. You can explore the full interim leaderboard on our Hugging Face space.

A few things to note:

We removed entries suspected of violating challenge rules.
We have been a little disappointed by a small number of participants who have been trying to circumvent our challenge rules. As a reminder, only one entry per team is allowed. Our system automatically overwrites multiple entries with the same username, keeping only the latest one.
Using alternative accounts to circumvent this rule is not allowed, and we reserve the right to remove entries and block users that we suspect are engaging in these actions. If you believe your entry was omitted by mistake, reach out to us via email.

We use a tier-based system to group entries that are not significantly different from each other.
To account for statistical noise, we group submissions that perform similarly into tiers using the Benjamini-Hochberg procedure (detailed in our previous blog post). Starting with the top entry, subsequent submissions are grouped into Tier 1 until we encounter a model whose performance differs significantly, which is then assigned to Tier 2. In the Direct Inhibition track, only one entry made it into Tier 1, with the second entry being significantly different (Tier 2). For TDI, a whopping 26 entries were assigned to Tier 1 and are therefore indistinguishable according to this test. Below, we show the top 18 entries from the Direct Inhibition leaderboard (spanning Tiers 1–4) and all 26 Tier 1 entries from the TDI track..

The interim leaderboard might be a little different than the live leaderboard.
Because the live leaderboard reflects only half the test set, your rank may have moved up or down. The interim leaderboard provides a much clearer picture of what to expect in the final standings, so use this benchmark to fine-tune your approach for the home stretch!

Direct Inhibition Interim Leaderboard

Rank Username Significance Model report Proprietary data? Open-source code? MA-ST-RAE MA-MAE
1 preheat-to-450 Tier 1 NA ✘ ✘ 0.3721 0.6005
2 wbot Tier 2 NA ✘ ✘ 0.3912 0.6195
3 cypadeedoodah Tier 2 NA ✘ ✘ 0.3921 0.6511
4 Rubyy Tier 2 NA ✘ ✘ 0.4014 0.6394
5 EleanorRigby Tier 3 NA ✘ ✘ 0.4173 0.6627
6 beast Tier 3 NA ✘ ✘ 0.4263 0.6947
7 leebrubeck Tier 4 NA ✘ ✘ 0.4336 0.6264
8 hfger Tier 4 NA ✘ ✘ 0.4357 0.6383
9 first-chair Tier 4 NA ✘ ✘ 0.4362 0.6306
10 furtivepeony Tier 4 NA ✘ ✘ 0.4370 0.6383
11 Scigantic Tier 4 NA ✘ ✘ 0.4382 0.6568
12 psk Tier 4 NA ✔ ✘ 0.4387 0.6424
13 hf-far Tier 4 NA ✘ ✘ 0.4410 0.6680
14 rasayan-labs Tier 4 link ✘ ✘ 0.4413 0.6449
15 briford Tier 4 link ✘ ✔ 0.4430 0.6415
16 horse-pusher Tier 4 NA ✘ ✘ 0.4440 0.6689
17 as-ris-1 Tier 4 NA ✘ ✘ 0.4444 0.6425
18 curiousone Tier 4 NA ✘ ✘ 0.4469 0.6447

TDI Interim Leaderboard

Rank Username Significance Model report Proprietary data? Open-source code? MA-MCC MA-Accuracy
1 nova Tier 1 NA ✘ ✘ 0.3980 0.8552
2 CYPblend Tier 1 NA ✘ ✘ 0.3888 0.8387
3 jeremy Tier 1 link ✘ ✔ 0.3868 0.8451
4 TeamPozeSCAF Tier 1 NA ✘ ✘ 0.3847 0.8544
5 N283T Tier 1 NA ✘ ✘ 0.3798 0.8478
6 stir_bar Tier 1 link ✘ ✘ 0.3796 0.8397
7 harinu123 Tier 1 NA ✘ ✘ 0.3767 0.8464
8 ADI Tier 1 NA ✘ ✘ 0.3732 0.8505
9 Asidsal11 Tier 1 NA ✘ ✘ 0.3729 0.8591
10 TeamPrescience Tier 1 NA ✘ ✘ 0.3714 0.8502
11 preheat-to-450 Tier 1 NA ✘ ✘ 0.3653 0.8357
12 lex Tier 1 NA ✘ ✘ 0.3622 0.8409
13 PIPL-Bio Tier 1 NA ✘ ✘ 0.3609 0.8389
14 EleanorRigby Tier 1 NA ✘ ✘ 0.3607 0.8362
15 newcyp Tier 1 NA ✘ ✘ 0.3558 0.8162
16 NIPERKOLKATA_PI_LAB Tier 1 NA ✘ ✘ 0.3553 0.8364
17 NorthCoil Tier 1 NA ✘ ✘ 0.3553 0.8113
18 C_CYPher-graphattn-B2regaux Tier 1 NA ✘ ✘ 0.3543 0.8065
19 curiousone Tier 1 NA ✘ ✘ 0.3533 0.8349
20 horse-pusher Tier 1 NA ✘ ✘ 0.3522 0.8337
21 yourchoice Tier 1 NA ✘ ✘ 0.3468 0.8246
22 450nm Tier 1 link ✘ ✘ 0.3457 0.8376
23 Tutu Tier 1 NA ✘ ✘ 0.3451 0.8173
24 furtivepeony Tier 1 NA ✘ ✘ 0.3443 0.8280
25 srijitseal Tier 1 NA ✘ ✘ 0.3437 0.7773
26 aries Tier 1 NA ✘ ✘ 0.3406 0.8400

The community has risen up to the challenge

The response from the community has been phenomenal! We’ve seen steady growth since the start of the challenge, with over 200 unique participants in the Direct Inhibition (Regression) track and over 100 in the TDI (Classification) track.

But the challenge is far from over. If you have been waiting to participate, now is the time to jump in! You still have over a month to submit your predictions before the challenge closes on November 3, 2026.

Time to work on your CYP2D6!

Here’s the trend we have seen so far: CYP3A4 has been the easiest isoform to predict across both tracks, while CYP2D6 has proven much trickier (see figure below).

To find why, we need to look at how the dataset was constructed. As detailed in our announcement blog post, the test set compounds were selected from a hit expansion on the most potent compounds on the training set. For three of the isoforms (namely, CYP3A4, CYP1A2, and CYP2C9), we took the top 25 hits and expanded each into 10 chemo-similar analogues (750 compounds total). Because CYP2D6 hits were not explicitly prioritized in this expansion, the test set for CYP2D6 contains inherently less potent compounds than the other three isoforms. This distribution shift between training and test sets makes CYP2D6 a much harder target to predict in this competition.

We’re only halfway through the challenge, so we can’t wait to see what creative strategies you come up with to boost performance on CYP2D6 over the coming weeks!

And now (*drumroll*) we introduce our new track

Today we also mark the start of a new Structure Track which will run alongside the inhibition tracks through the end of the competition.

Our team at UCSF has resolved 20 novel Cryo-EM structures of the CYP3A4 isoform complexed with a selection of assayed ligands (15 from the training set, 2 from the test set, and 3 additional compounds). Participants are tasked with predicting the 3D structure of a protein-ligand complex given only the SMILES strings.

CYP3A4 features a remarkably flexible binding pocket that undergoes significant conformational changes to accommodate diverse chemical matter. This elasticity, combined with the fact that similar ligands can adopt drastically distinct binding poses, makes CYP3A4 a nightmare for computational structure prediction. These 20 high-quality Cryo-EM structures significantly expand the chemical space beyond existing X-ray PDB structures. For more details on the experimental design, check out our recent preprint.

How to participate

You will submit predicted 3D structures for all 20 complexes compiled into a single .zip archive. Feel free to leverage co-folding models, docking software, public PDB structures, or proprietary data. For inspiration, we recommend looking into some of the workflows that were successful in our previous PXR challenge, but we are also looking forward to seeing all the creative approaches y’all may come up with!

Full instructions for loading data and formatting submissions are available in our tutorial.

Rules

Submissions will be scored primarily using the Local Difference Test for Protein Ligand Interactions (LDDT-PLI), while other metrics ,such as Binding-Site Superposed, Symmetry-Corrected Pose Root Mean Square Deviation (BiSyRSMD), will be provided as secondary metrics. Comparisons are run automatically against the blinded ground-truth using the OpenStructure pipeline. Any structure that fails basic OpenStructure checks receives an LDDT-PLI score of 0 and a BiSyRMSD penalty of 20 Å.

What’s new this time? We’re integrating PoseBusters into the scoring framework. Co-folding algorithms often optimize for raw atom placement at the expense of physical plausibility (e.g., clashing atoms or strained geometries). This was demonstrated by Buttenschoen, et. al., who designed a set of tests to validate chemical and geometric consistency of a ligand pose. We believe that accuracy metrics, such as LDDT-PLI or BiSyRSMD don’t tell the whole story ,and producing realistic poses is equally important, so every structure will now pass through PoseBusters physical/chemical sanity checks. If a prediction fails these checks, it will receive an LDDT-PLI score of 0. Look out for an upcoming blog post where we run a retrospective PoseBusters analysis on our earlier PXR challenge structures!

As with the other tracks, submission feedback will be given automatically in a dedicated Discord channel, including a list of structures that don’t pass either OpenStructure or PoseBuster checks. In addition, all the scripts we use for structure evaluation are provided in the tutorial repo, so you can evaluate your models locally prior to submission.

On to the last dance

The platform setup remains unchanged for the final phase. The live leaderboard will stay active in our Hugging Face Space until the final deadline: The platform setup remains unchanged for the final phase. The live leaderboard will stay active in our Hugging Face Space until the final deadline: November 3, 2026, at 11:59 PM UTC.

A quick checklist as we approach the end of the challenge:

  • Use the interim leaderboard results to adjust your workflows for your final submission. Keep in mind that performance on the live leaderboard is not a true reflection of your performance on the final standings
  • Remember that we require a comprehensive model report along with your final submission. If this report is not included by the time the submission closes, you will not be considered for the final leaderboard.
  • Finally, we maintain a strict policy against multiple submissions per team; any attempts to circumvent this policy may result in disqualification at our own discretion

Good luck and happy modeling!

Acknowledgements

We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute), Schrödinger Inc, and the Gates Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.

This work is supported by the Advanced Research Projects Agency for Health (ARPA-H) under AVOID-OME, and Award Number 1AY1AX000035. The contents are those of the authors. They may not reflect the policies of the Department of Health and Human Services or the U.S. government. The content is solely the responsibility of the authors and does not necessarily represent the official views of the Advanced Research Projects Agency for Health.