Related Experiment Video
Updated: Feb 24, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Developing Large Language Model-based Pipeline for Identification of Disease Diagnosis: A Case Study on Identifying
Mei Wang1,2, Yuan-Hung Kuan1,2, Patrik R Alba3,4
1Research Service, St. Louis Veterans Affairs Medical Center, St. Louis, MO.
Abstract:
Accurately identifying disease diagnoses from electronic health records (EHRs) is crucial for clinical/biomedical research; however, this is challenging when diagnoses are complex and require data from several sources, e.g., multiple myeloma (MM) and its precursor condition, MGUS. Leveraging the national Veterans Health Administration EHRs, we developed and validated a large language model (LLM)-based pipeline that utilizes only clinical notes from randomly selected patients identified via ICD codes for MGUS/MM. Among the evaluated LLMs and alternative approaches, Llama-3-8B-based pipeline with prompt engineering achieved the best performance. This pipeline not only saved the preprocessing steps and shortened the overall processing time but also outperformed rule-based or machine learning-based methods for identifying MGUS and achieved comparable performance for MM, solely relying on clinical notes. Our work demonstrates that the developed LLM-based pipeline can efficiently and effectively identify MGUS/MM diagnoses to replace manual chart abstraction and rule- or machine learning-based natural language processing methods.
