Practicing What We Preach: Why Open Science Needs Clear Rules
At OpenADMET, open science isn't just a buzzword for us. We're tearing down silos by sharing raw data, open-source models, and runnable code instead of static papers. Through blind challenges and real-time collaboration, we're building a faster, transparent future for drug discovery.
Author: Pat Walters
DOI: Pending
At OpenADMET, "Open" is our first name. It is not just a moniker, but our core operating philosophy. We believe the future of ADMET modeling depends on breaking down data silos, promoting radical transparency, and fostering an active, collaborative global community. However, advocating for open science is easy; building an organization that consistently lives up to those ideals across every dataset, model, and publication requires deliberate choices and constant accountability. Practicing what we preach means setting clear standards for how we work and holding ourselves to the same bar we hold others to.
I am, and have always been, an open science advocate. I joined OpenADMET to devote myself full-time to open science, building on a career shaped by years of maintaining open-source code, writing a Practical Cheminformatics blog, and sharing data openly. Over the past 25 years, I have seen firsthand the transformative impact that community benchmarks and blind challenges, including CASP, CACHE, D3R, and SAMPL, can have on our field. They offer an objective, fast, and rigorous way to do science that keeps pace with modern machine learning. On a personal level, nothing is more rewarding than seeing hundreds of researchers from around the world gather on the OpenADMET Discord server to troubleshoot code, challenge assumptions, and learn from one another in real time.
To turn these values into concrete operational reality, we established the OpenADMET Core Principles & Community Policies. These commitments address the bottlenecks of legacy scientific infrastructure, which was simply not designed for the current pace of AI and ML innovation:
- Speed and Executable Science Over Static Papers: Traditional journal review cycles take months, while ML models and methods evolve weekly. Our policy prioritizes immediate communication through preprints and technical blog posts, paired directly with runnable code, Jupyter/Marimo notebooks, and open GitHub repositories. Science should be accessible and executable from day one, not locked behind paywalls or trapped in static PDFs.
- Radical Data Transparency from First Principles: Machine learning models are only as trustworthy as the data beneath them. Under our policies, all chemical structures must be fully disclosed as standardized, machine-readable SMILES, with no obscured or proprietary scaffolds. We also commit to releasing instrument-level raw data, such as LC/MS traces and plate reader outputs, and depositing datasets in public repositories such as ChEMBL and the PDB, so that anyone can verify and re-curate the data from scratch.
- Open Models and Full Reproducibility: We build our models openly under permissive licenses such as MIT or Apache 2.0. That includes releasing model weights, training pipelines, data splits, and model cards that candidly report where an architecture fails.
- Community First Through Blind Challenges: Data is far more valuable in the hands of the community than locked in proprietary silos. Our blind challenges use fully public scoring scripts, transparent benchmark datasets, and clear rules, with all data released openly once a challenge ends.
Changing how an entire scientific discipline communicates and shares work is an ongoing, iterative process. We treat our policies as a living commitment: when we encounter new challenges or fall short, we will say so openly, explain what we are changing, and invite the community to help us improve.
We would like to thank our funders for their support of OpenADMET, in particular ARPA-H, Radial (part of the Astera Institute (https://ror.org/00ydx1s47)), Schrödinger Inc., the Gates Foundation, and the OpenAI Foundation. We would also like to thank our partners Enamine, HuggingFace, OpenEye, CDD Vault, Discovery Life Sciences, and the beamline staff at NSLS-II for their support.
Finally, we are deeply grateful to the hundreds of participants who bring energy and rigor to our challenges and community every day. Our goal remains the same: building a faster, fully transparent, and community-driven future for predictive science.