Related Experiment Video
Updated: Jan 16, 2026

Biosensor-based High Throughput Biopanning and Bioinformatics Analysis Strategy for the Global Validation of Drug-protein Interactions
Published on: December 1, 2020
Enabling Open Machine Learning of Deoxyribonucleic Acid-Encoded Library Selections to Accelerate the Discovery of
James Wellnitz1, Shabbir Ahmad2, Nabin Bagale3
1Division of Chemical Biology and Medicinal Chemistry, UNC Eshelman School of Pharmacy, University of North Carolina, Chapel Hill, North Carolina 27516, United States.
Abstract:
Machine learning (ML) is increasingly used in DNA-encoded library (DEL) screening for ligand discovery, but its success depends on access to suitable data sets, which are often proprietary and costly. To overcome this, we present the first fully open, automated DEL-ML framework using public DEL data sets and chemical fingerprints to enable reproducible, accessible drug discovery. Our workflow─from model training to virtual screening and compound selection─requires no human intervention. As a proof of concept, we identified binders for WDR91 by training ML models on the HitGen OpenDEL library (3B molecules) and screening the Enamine REAL Space library (37B molecules), yielding 50 candidates. Experimental testing confirmed seven novel binders with dissociation constants between 2.7-21 μM. Our open-source approach matches the performance of proprietary methods, demonstrating that public DEL data can support robust ML-driven ligand discovery and fostering transparency and broader community participation in drug development.

