Related Experiment Video
Updated: Feb 14, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Integrating artificial intelligence and manual curation to enhance bioassay annotations in ChEMBL
Ines Smit1, Melissa F Adasme1, Emma Manners1
1European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridgeshire, CB101SD, UK.
Enhancing ChEMBL bioassay data quality is crucial for cheminformatics and machine learning (ML). Recent efforts improve assay metadata standardization and machine readability, boosting FAIR data principles for robust compound-target modeling.
Area of Science:
- Cheminformatics
- Bioactivity data analysis
- Machine learning applications
Background:
- The ChEMBL database contains a growing volume and diversity of bioactivity data.
- Standardized, interoperable, and machine-readable assay metadata is critical for cheminformatics and machine learning (ML) applications.
- Current ChEMBL bioassay annotations require enhancement for improved data quality and granularity.
Purpose of the Study:
- To present recent efforts to enhance the quality and granularity of bioassay annotations in the ChEMBL database.
- To improve the standardization, interoperability, and machine readability of assay metadata.
- To enable more robust downstream analyses and precise compound-target activity modeling.
Main Methods:
- Implemented a "perfect assay description" template for consistent annotation.
- Utilized natural language processing (NLP) techniques and multi-class classification for automated extraction of assay parameters and categorization of legacy data.
- Developed and validated a spaCy-based Named Entity Recognition (NER) model to identify experimental methods.
- Created a complementary classification model to refine ASSAY_TYPE categorization.
- Improved metadata extraction for ADME endpoints, organism and protein variant annotations, and ontology linking using tools like text2term.
Main Results:
- The "perfect assay description" template guides consistent annotation practices.
- NLP and classification models automatically extract key assay parameters and assign broad assay categories.
- The spaCy-based NER model achieves high precision and recall in identifying experimental methods.
- A complementary classification model enhances ASSAY_TYPE categorization beyond the existing schema.
- Improvements in metadata extraction and ontology linking advance the FAIRness of ChEMBL bioassay data.
Conclusions:
- Recent enhancements significantly improve the quality and granularity of ChEMBL bioassay annotations.
- These improvements advance the Findable, Accessible, Interoperable, and Reusable (FAIR) principles for ChEMBL bioassay data.
- The enhanced data enables more robust downstream analyses and more precise compound-target activity modeling in cheminformatics and ML applications.
Related Concept Videos
Genome Annotation and Assembly
Intelligence
Measures of Intelligence
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Multiple Intelligences Theory
Cattell's Theory of Intelligence
Fluid intelligence involves the capacity to solve new problems and adapt to unfamiliar situations. It's the type of intelligence individuals use when they encounter a novel problem or puzzle that requires innovative thinking. For instance, figuring out how to operate a new gadget relies heavily on...
Triarchic Theory of Intelligence

