Related Experiment Video
Updated: May 5, 2026

Detecting Somatic Genetic Alterations in Tumor Specimens by Exon Capture and Massively Parallel Sequencing
Published on: October 19, 2013
Benchmarking Generative AI Tools for Interpretation of the WHO TB Mutation Catalogue
Miguel Moreno-Molina1, Anita Suresh1, Rebecca E Colman1,2
1Foundation for Innovative New Diagnostics (FIND), Geneva, Switzerland.
Abstract:
The World Health Organization (WHO) 2023 Mutation Catalogue for Mycobacterium tuberculosis is a crucial knowledgebase and tool for clinical interpretation of mutations associated with drug-resistant TB. However, the document's complexity and size pose challenges for many users. This study evaluated the potential of generative artificial intelligence (AI) models to facilitate natural language user interaction with the catalogue. This was a benchmarking study, not a clinical usability trial. Four prominent AI models-Google Gemini 2.5 Pro, OpenAI ChatGPT 4.1, Perplexity AI, and DeepSeek R1-were assessed through general test questions, mutation search and retrieval tasks using both full catalogue queries and antibiotic-specific tables, and the application of additional grading rules to score novel mutations. Performance was measured based on accuracy, completeness, clarity, source citation, and the presence of hallucinations. Google Gemini 2.5 Pro consistently demonstrated superior performance in accuracy, completeness, and avoidance of hallucinations across most evaluations, especially in general queries and large dataset searches. DeepSeek R1 excelled in applying grading rules to novel mutations and showed high accuracy in focused datasets, but exhibited some hallucinations. ChatGPT 4.1 was strong in clarity but lacked proper citations, and Perplexity AI showed variable performance with a higher frequency of hallucinations. The findings highlight the potential of AI tools to enhance the accessibility of complex knowledgebases like the WHO Mutation Catalogue, while emphasizing the need for rigorous benchmarking. While no model is yet suitable for direct clinical use, the results suggest that with further development, models like Google Gemini 2.5 Pro could form the basis of a custom AI agent to assist users in navigating this critical resource, ultimately contributing to improved TB control efforts.
More Related Videos
08:46Implementation of In Vitro Drug Resistance Assays: Maximizing the Potential for Uncovering Clinically Relevant Resistance Mechanisms
Published on: December 9, 2015
09:40Author Spotlight: Unveiling the Role of TMOD3 in Platinum Resistance and Immune Infiltration in Ovarian Cancer
Published on: August 2, 2024