Related Experiment Video
Updated: Jun 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
From knowledge generation to knowledge verification: examining the biomedical generative capabilities of ChatGPT
Ahmed Abdeen Hamed1,2,3, Alessandro Crimi4, Magdalena M Misiak5
1MGEN - College of Engineering, Northeastern University Miami, Miami, FL 33127, USA.
Large language models (LLMs) can accelerate biomedical research but require validation. This study developed a method to computationally verify LLM-generated disease associations, finding high accuracy for terms but variable reliability for IDs, suggesting careful use and Retrieval Augmented Generation (RAG) for enhanced trust.
Area of Science:
- Computational biology
- Biomedical informatics
- Artificial intelligence in healthcare
Background:
- Large language models (LLMs) offer potential for accelerating scientific tasks.
- Concerns exist regarding the factual authenticity and reliability of LLM-generated biomedical knowledge.
- Need for robust methods to evaluate the accuracy of AI-generated biomedical information.
Purpose of the Study:
- To present a computational approach for evaluating the factual accuracy of biomedical knowledge generated by LLMs.
- To assess the consistency and reliability of disease-centric associations (drugs, symptoms, genes) generated by various ChatGPT models.
- To identify limitations and potential improvements for using LLMs in biomedical knowledge discovery.
Main Methods:
- Developed a computational framework to generate disease-centric associations using prompt-engineered LLMs (ChatGPT, GPT-4, GPT-4o, GPT-4o-mini).
- Verified generated associations against established biomedical ontologies.
- Assessed accuracy for identifying disease, drug, and genetic terms, and coverage of disease-drug, disease-gene, and disease-symptom relationships.
Main Results:
- High accuracy in identifying disease terms (88%-97%), drug names (90%-91%), and genetic information (88%-98%).
- Lower accuracy for symptom term identification (49%-61%) attributed to informal descriptions.
- Good coverage for disease-drug (89%-91%) and disease-gene (89%-91%) associations, with lower coverage for disease-symptom associations (49%-62%).
- Generated identifiers were frequently invalid or redundant despite high term accuracy.
Conclusions:
- LLM-generated biomedical knowledge can be reliable when used cautiously.
- Prompt engineering and verification against ontologies are crucial for accuracy.
- Retrieval Augmented Generation (RAG) shows promise for enhancing the reliability of LLM outputs in biomedical applications.
Related Concept Videos
What is Genetic Engineering?
Non-equilibrium in the Cell
Transgenic Plants
The first-ever transgenic plant was a tobacco plant developed in 1983 that showed resistance against the tobacco mosaic virus. Since then, many transgenic plants have been developed and commercialized for improving the agricultural, ornamental, and horticultural value of a crop plant. Transgenic...

