Related Experiment Video
Updated: Jul 8, 2025

Implementation of In Vitro Drug Resistance Assays: Maximizing the Potential for Uncovering Clinically Relevant Resistance Mechanisms
Published on: December 9, 2015
Asking the right questions for mutagenicity prediction from BioMedical text
Sathwik Acharya1, Nicolas K Shinada1,2, Naoki Koyama3
1The Systems Biology Institute, Tokyo, Japan.
This study introduces a new method to predict chemical mutagenicity from scientific texts, aiding drug development. The MutaPredBERT model automates the extraction of mutagenicity data, improving knowledge base creation.
Area of Science:
- Computational chemistry
- Bioinformatics
- Drug discovery
Background:
- Assessing chemical mutagenicity is crucial for drug development.
- Existing mutagenicity databases require laborious manual curation and are difficult to update.
- Automating the extraction of mutagenicity data from scientific literature is needed.
Purpose of the Study:
- To propose and address the problem of predicting chemical mutagenicity directly from textual information in scientific publications.
- To develop a model that can predict mutagenicity based on natural language descriptions of chemical evidence.
- To establish the utility of large language models for constructing structured knowledge bases from unstructured scientific text.
Main Methods:
- Construction of a gold-standard dataset for chemical mutagenicity.
- Development of MutaPredBERT, a prediction model fine-tuned on BioLinkBERT.
- Formulation of the prediction task as a question-answering problem.
- Leveraging transfer learning with large transformer-based models.
Main Results:
- Achieved a Macro F1 score greater than 0.88.
- Demonstrated high prediction accuracy even with a relatively small fine-tuning dataset.
- Successfully applied a question-answering approach to mutagenicity prediction.
Conclusions:
- The MutaPredBERT model effectively predicts chemical mutagenicity from scientific literature.
- Large language models are valuable tools for automating the creation of structured knowledge bases.
- This approach offers a more practical and scalable alternative to manual database curation for mutagenicity information.
Related Concept Videos
Mutagenicity and Carcinogenicity
Mismatch Repair
The Mutator Protein Family Plays a Key Role in DNA Mismatch Repair
The human genome has more than 3 billion base pairs of DNA per cell. Prior to cell division, that vast amount of genetic...
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...
Complementation Tests
Organisms heterozygous for different mutations are crossed pairwise in all combinations. If present on different genes, the mutations can complement each other by providing the missing...
In-vitro Mutagenesis
Cancer-Critical Genes II: Tumor Suppressor Genes
When the function of certain critical genes, especially those involved in cell cycle regulation and cell growth signaling cascades, gets disrupted, it upsets the cell cycle progression. Such cells with unchecked cell cycles start proliferating uncontrollably and eventually develop into tumors.
Such genes that act...

