Related Experiment Video
Updated: Sep 26, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Hierarchical prompting with reasoning large language models for immune checkpoint inhibitor-associated myocarditis
Enshuo Hsu1,2, Sheng-Chieh Lu3, Sara Ebrahimi3
1Enterprise Development and Integration, University of Texas MD Anderson Cancer Center, Houston, TX, USA.
Aims:
Immunotherapy with immune checkpoint inhibitors (ICIs) is an effective treatment for many cancers, but it can induce immune-related adverse events (irAEs). Among those, myocarditis is a severe complication associated with high morbidity and mortality. To support the rapidly emerging research, labour-intensive manual data extraction is often needed. Recent studies adopted large language models (LLMs) to systematically process clinical notes and identify patients with irAEs. However, research gaps persist: (i) For ICI-associated myocarditis, the applicability and scalability of LLMs have not been systematically evaluated; (ii) most existing studies focused on patient-level irAE detection while omitting the clinical nuances; and (iii) cutting-edge LLM techniques have not been adopted. Our aim was to evaluate an LLM approach to extract ICI-associated myocarditis evidence, including inflammatory infiltrate, life-threatening arrhythmias, heart failure, and stroke, from clinical notes.
Methods And Results:
We collected clinical notes of patients with ICI-associated myocarditis at a single centre to develop and evaluate LLM-based methods that convert free-text clinical notes into structured data with entities (e.g. diagnosis, treatment, and imaging) and attributes (e.g. date, assertion, and status). We systematically evaluated three techniques: LLM reasoning, context engineering, and hierarchical prompting against ground truth created by a cardio-oncologist. Our proposed sentence-based context engineering and hierarchical prompting method, with a reasoning LLM, is significantly more accurate than the research trainees (F1 score of 0.6332 vs. 0.4933, P < 0.001), faster (37 min vs. 72 h per 1000 notes), and more cost-effective ($2.04 USD per 1000 notes).
Conclusion:
We proposed an optimized method that integrates hierarchical prompting, context engineering, and a reasoning LLM for ICI-associated myocarditis evidence extraction to support physicians and researchers in rapid clinical data collection for research to make diagnostic and treatment inferences.