Related Experiment Video
Updated: Oct 20, 2025

Mass Spectrometry-Guided Genome Mining as a Tool to Uncover Novel Natural Products
Published on: March 12, 2020
Models and Processes to Extract Drug-like Molecules From Natural Language Text
Zhi Hong1, J Gregory Pauloski1, Logan Ward2
1Department of Computer Science, University of Chicago, Chicago, IL, United States.
This study introduces a hybrid AI and human approach to identify potential drug molecules from scientific literature for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) research. The method efficiently extracts drug candidates, accelerating drug discovery efforts.
Area of Science:
- Computational Biology
- Bioinformatics
- Natural Language Processing
Background:
- Drug repurposing and discovery are critical for combating SARS-CoV-2.
- Scientific literature contains vast information on potential drug-like molecules.
- Existing Named Entity Recognition (NER) models struggle with specialized vocabularies in viral research literature.
Purpose of the Study:
- To develop and evaluate a human-AI hybrid pipeline for identifying drug-like molecules from scientific text.
- To address the limitations of existing NER models in processing large, specialized scientific corpora.
- To accelerate the identification of potential drug candidates for SARS-CoV-2 research.
Main Methods:
- Implemented an iterative model-in-the-loop strategy for efficient training data generation.
- Utilized scarce human expertise judiciously to train a specialized NER model.
- Applied the developed NER model to the COVID-19 Open Research Dataset Challenge (CORD-19) corpus.
Main Results:
- Achieved an F-1 score of 80.5% for the NER model, comparable to human performance.
- Identified 10,912 putative drug-like molecules within the CORD-19 corpus.
- Enriched computational screening targets by 3,591 molecules, with 18 ranking in the top 0.1% for 3CLPro docking.
Conclusions:
- The human-AI hybrid pipeline effectively identifies drug-like molecules from large scientific datasets.
- This approach significantly enhances the efficiency of drug discovery for viral diseases like COVID-19.
- The method offers a scalable solution for extracting valuable molecular information from scientific literature.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
Drug Discovery: Overview
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Drug Biotransformation: Overview
Drug Nomenclature
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Targets for Drug Action: Overview
Receptors are either membrane-spanning or intracellular proteins, which upon binding a ligand, get activated and transmit the signal downstream to elicit a response. Drugs bind receptors, either mimicking the action of endogenous ligands or blocking the receptor activity to bring about a modified response. Nearly 35% of approved drugs target the G...