Related Experiment Video
Updated: Jun 13, 2026

Identification of Circular RNAs using RNA Sequencing
Published on: November 14, 2019
Benchmarking large language models for cell-free RNA diagnostic biomarker discovery.
Hunter A Gaudio1, Andrew Bliss1, Conor J Loy1
1Meinig School of Biomedical Engineering, Cornell University, Ithaca, NY, USA.
Large language models show promise for biomarker discovery from omics data. While facing challenges, they can identify diagnostic gene panels and automate classifier construction for various diseases.
Area of Science:
- Biomedical Informatics
- Artificial Intelligence in Medicine
- Genomics and Computational Biology
Background:
- Large language models (LLMs) offer potential for analyzing complex biomedical data, including high-throughput omics, for biomarker discovery.
- Biomarker identification is crucial for diagnosing and managing diseases like Kawasaki disease, tuberculosis, and myalgic encephalomyelitis/chronic fatigue syndrome.
Purpose of the Study:
- To benchmark the performance of six leading large language models in biomarker discovery from plasma cell-free RNA data.
- To evaluate LLMs for both literature-guided gene panel nomination and autonomous end-to-end diagnostic classifier construction.
- To assess the capabilities and limitations of current LLMs in clinical diagnostics.
Main Methods:
- Six LLMs (OpenAI, Anthropic, Google) were evaluated on plasma cell-free RNA datasets from three distinct clinical cohorts.
- Performance was assessed for literature-guided nomination of diagnostic gene panels for machine learning.
- Autonomous construction of end-to-end classifiers from raw omics data to predictions was also evaluated.
Main Results:
- LLM-nominated gene panels, despite prompt adherence issues, identified relevant immune pathways and outperformed random panels.
- LLM performance matched differential gene expression baselines in the tuberculosis cohort.
- End-to-end automation feasibility varied by model and task; one LLM achieved near-conventional performance for one cohort, but performance declined for others.
Conclusions:
- LLMs demonstrate potential in biomarker discovery and diagnostic classifier development from omics data, though prompt adherence and task dependency are limitations.
- LLM-nominated gene panels can capture biological relevance and serve as a viable alternative to traditional methods.
- Further development is needed to optimize LLMs for robust and reliable clinical diagnostic applications.
More Related Videos
12:54Real-time Analysis of Transcription Factor Binding, Transcription, Translation, and Turnover to Display Global Events During Cellular Activation
Published on: March 7, 2018
12:44Identification of Key Factors Regulating Self-renewal and Differentiation in EML Hematopoietic Precursor Cells by RNA-sequencing Analysis
Published on: November 11, 2014