Making large language models reliable data science programming copilots for biomedical research
Zifeng Wang1,2, Benjamin Danek1,2, Ziwei Yang3
1Keiji AI, Seattle, WA, USA.
None:
Large language models (LLMs) can generate impressive data visualizations from simple requests, yet their accuracy remains underexplored. Here we present a benchmark of 293 coding tasks derived from 39 studies across 7 biomedical research areas, including biomarkers, integrative analysis, genomic profiling, molecular characterization, therapeutic response, translational research and pan-cancer analysis. Benchmarking eight proprietary and eight open-source LLMs under various prompting strategies reveals an overall accuracy below 40%. This low accuracy raises serious concerns about the risk of propagating incorrect scientific findings when blindly relying on AI-generated analyses. Therefore, we develop an AI agent that begins with and iteratively refines an analysis plan before generating code, achieving 74% accuracy. We embody this insight in a platform that enables users to codevelop analysis plans with LLMs and execute them within an integrated environment. In a user study with five medical researchers, the platform enabled users to complete over 80% of the analysis code for three studies.
More Related Videos
Related Concept Videos
Reliability and Validity
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Psychology as a Science
The scientific method in psychology involves six critical steps: making observations, formulating hypotheses, conducting tests, analyzing...
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Language and Cognition


