Related Experiment Video
Updated: Sep 24, 2025

Rare Event Detection Using Error-corrected DNA and RNA Sequencing
Published on: August 3, 2018
Machine Learning on DNA-Encoded Library Count Data Using an Uncertainty-Aware Probabilistic Loss Function.
Katherine S Lim1,2, Andrew G Reidenbach3, Bruce K Hua3,4
1Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, 77 Massachusetts Avenue, Cambridge, Massachusetts 02139, United States.
This study introduces a novel regression approach for DNA-encoded library (DEL) screening data, improving the analysis of molecular enrichments and enabling better identification of drug candidates. The method effectively denoises data and visualizes structure-activity relationships for drug discovery.
Area of Science:
- Computational chemistry
- Cheminformatics
- Drug discovery
Background:
- DNA-encoded library (DEL) screening and quantitative structure-activity relationship (QSAR) modeling are key techniques in drug discovery.
- Current QSAR modeling of DEL data often uses binary classifiers on aggregated 'disynthons', which can lose information and fail to capture enrichment levels.
Purpose of the Study:
- To develop a regression approach for analyzing DEL selection data that overcomes limitations of binary classification.
- To enable more nuanced understanding of molecular enrichments and facilitate the identification of novel drug candidates.
Main Methods:
- Developed a regression model to learn DEL enrichments of individual molecules.
- Implemented a custom negative-log-likelihood loss function to denoise sparse and noisy DEL data.
- Modeled the Poisson statistics of the sequencing process inherent in DEL workflows.
Main Results:
- The regression approach effectively denoises DEL data and visualizes structure-activity relationships.
- Models can identify low-confidence outliers by accounting for data uncertainty.
- Demonstrated on large DEL datasets screened against carbonic anhydrase (CAIX), soluble epoxide hydrolase (sEH), and SIRT2.
Conclusions:
- The proposed regression method offers a powerful tool for analyzing DEL data, improving the identification of structure-activity trends and enriched pharmacophores.
- This uncertainty-aware regression approach is applicable to other sparse or noisy datasets with known stochasticity.
- Enhances the utility of DEL screening in drug discovery pipelines.
Related Concept Videos
Propagation of Uncertainty from Random Error
Propagation of Uncertainty from Systematic Error
Uncertainty: Overview
DNA Microarrays
Uncertainty: Confidence Intervals
Genome Copying Errors

