Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Supervised Fine-Tuning of Large Language Models With Chain-of-Thought Reasoning for Pediatric Heart Disease Detection

Haoming Shi1,2,3, Justin B Long4,5, Michael C Fiedorek4,5

  • 1Department of Biomedical Engineering, Duke University, 1427 Fitzpatrick Center Box 90281, Durham, NC, 27708, United States, 1 9196695131.

JMIR Formative Research
|June 8, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

A scoping review of the evidence supporting antiarrhythmic use after paediatric cardiac surgery.

Cardiology in the young·2026
Same author

Clinical Criteria for the Definition of Refractory Septic Shock: A Joint Delphi Consensus from the Society of Critical Care Medicine (SCCM) and European Society of Intensive Care Medicine (ESICM).

Critical care medicine·2026
Same author

Clinical criteria for the definition of refractory septic shock: a joint Delphi consensus from the Society of Critical Care Medicine (SCCM) and European Society of Intensive Care Medicine (ESICM).

Intensive care medicine·2026
Same author

Predicting pulmonary hypertension in infants with bronchopulmonary dysplasia.

Journal of perinatology : official journal of the California Perinatal Association·2026
Same author

rECMOmender: Reinforcement Learning for Decision Support in Venovenous Extracorporeal Membrane Oxygenation Management.

Critical care explorations·2026
Same author

Continuous Physiologic Markers of Heart Rate Variability Derived From Bedside Electrocardiogram Precede Onset of Acute Respiratory Distress Syndrome: A Physiologic Modeling Study.

Critical care explorations·2025

Supervised fine-tuning of large language models (LLMs) effectively detects pediatric heart disease (PHD) in echocardiogram reports. Chain-of-thought (CoT) reasoning improved internal accuracy and consistency, though external validation showed varied results.

Area of Science:

  • Artificial Intelligence in Medicine
  • Natural Language Processing
  • Clinical Informatics

Background:

  • Pediatric heart disease (PHD) is often under-documented in electronic health records.
  • Extracting PHD from unstructured echocardiogram reports is challenging.
  • Automated methods can improve PHD data quality for clinical support and research.

Purpose of the Study:

  • To assess the feasibility of using supervised fine-tuning of large language models (LLMs) for PHD detection.
  • To compare LLMs with and without chain-of-thought (CoT) reasoning for characterizing PHD.
  • To analyze unstructured echocardiogram reports for clinically significant or historical PHD.

Main Methods:

  • Developed a PHD detection algorithm using fine-tuned open-source LLMs (LLaMA, Qwen).
Keywords:
congenital heart defectsechocardiogramlarge language modelmachine learningnatural language processing

Related Experiment Videos

Last Updated: Jun 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

  • Analyzed 9,749 echocardiogram reports, with 712 adjudicated by experts.
  • Incorporated R1-generated CoT reasoning into LLM prompts for fine-tuning.
  • Main Results:

    • The fine-tuned Qwen-7B model with CoT achieved the highest internal accuracy (92.4%).
    • CoT-enhanced models showed improved classification consistency internally.
    • External validation revealed higher accuracy for non-CoT models, but Qwen-7B CoT offered more balanced performance.

    Conclusions:

    • Supervised fine-tuning of LLMs with CoT is effective for automated PHD detection in clinical notes.
    • CoT reasoning shows promise for improving LLM performance in medical text analysis.
    • Further validation is needed for real-world clinical decision support integration.