Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jun 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Screening for Missed Opportunities for Diagnosis in the ED Using eTriggers and Large Language Models.

Clifford M Marks1, Sean Gibney2, Bryan Stenson2

  • 1Department of Emergency Medicine, Georgetown University, Washington, DC.

JAMA Network Open
|June 29, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The Sepsis ImmunoScore Predicts Sepsis, Mortality, and Deterioration Better than Clinical Scores and Widely Available Biomarkers.

Diagnostics (Basel, Switzerland)·2026
Same author

Does Early Opioid Use for Relief of Acute Abdominal Pain Increase Use of Abdominal CT or Prolong Emergency Department Length of Stay?

Emergency medicine international·2026
Same author

Effect of Cannabis Use on Revision Surgery After Lumbar Spine Fusion: A Systematic Review and Meta-Analysis.

Clinical spine surgery·2026
Same author

Large reasoning models as thinking machines for medicine.

Nature biomedical engineering·2026
Same author

Improving End-of-Life Screening in the Emergency Department With Collaborative Artificial Intelligence.

Annals of emergency medicine·2026
Same author

Multimodal characterization of articular cartilage degeneration in the humeral head using Raman spectroscopy, biomechanics, and imaging.

Osteoarthritis and cartilage·2026

Large language models (LLMs) show varied performance in identifying missed opportunities for diagnosis (MODs) within emergency department (ED) electronic triggers (eTriggers). Model evaluation is crucial for workflow integration, as LLMs offer different sensitivity-specificity trade-offs.

Area of Science:

  • Artificial Intelligence in Healthcare
  • Clinical Decision Support Systems
  • Emergency Medicine Quality Improvement

Background:

  • Emergency department (ED) quality reviews often utilize administrative electronic triggers (eTriggers) but have limited success in detecting missed opportunities for diagnosis (MODs).
  • Commercial large language models (LLMs) present a potential solution for screening MODs, but real-world evaluation data are scarce.

Purpose of the Study:

  • To evaluate the effectiveness of various LLMs in identifying MODs within ED eTrigger cohorts.
  • To compare the performance of different LLMs across distinct ED patient populations.

Main Methods:

  • A retrospective diagnostic study analyzed 2 eTrigger cohorts: ED discharge with return hospital admission within 72 hours and ED admission to the floor with intensive care unit (ICU) escalation within 24 hours.

Related Experiment Videos

Last Updated: Jun 30, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

  • 288 encounters were evaluated by six LLMs (Claude Sonnet 4, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 3 Pro, GPT-5, GPT-5 mini) and adjudicated by emergency physicians using the Safer Dx framework.
  • Main Results:

    • Across 288 encounters with 39 MODs (13.5%), LLM performance varied. Sensitivity ranged from 42.9% to 85.7%, specificity from 55.9% to 82.9%, and AUC from 0.65 to 0.73 in the 72-hour return cohort.
    • In the floor-to-ICU cohort, sensitivity ranged from 5.6% to 55.6%, specificity from 64.6% to 97.5%, and AUC from 0.57 to 0.82.
    • LLMs demonstrated similar discrimination but different sensitivity-specificity trade-offs, with Claude Sonnet 4 favoring sensitivity and GPT-5 mini favoring specificity.

    Conclusions:

    • LLM performance in identifying MODs differs based on the specific cohort and model evaluated.
    • Findings highlight the need for careful evaluation of LLMs within clinical workflows to understand their sensitivity-specificity trade-offs before implementation.
    • Reviewer-like concordance offers a different perspective on model behavior compared to discrimination metrics.