Related Experiment Videos
Multimodal artificial intelligence-based decision support for stroke classification in medical emergency calls using
Emil Iversen1, Hege Ihle-Hansen2, Kari Krizak Halle3
1Norwegian National Advisory Unit on Emergency Medical Communication, Haukeland University Hospital, Bergen, Norway; Department of Prehospital Emergency Medicine, Institute of Nursing and Health Promotion, Faculty of Health Sciences, Oslo Metropolitan University, Oslo, Norway; Department of Clinical Medicine, University of Bergen, Bergen, Norway.
Background:
Identifying stroke during medical emergency calls is challenging because of unstructured information, communication barriers, and heterogeneous symptoms. Artificial intelligence (AI) may improve early stroke recognition;however,evidence from emergency call settings remains limited, and the value of combining call transcripts with structured clinical data is unclear.We developed and internally validated a multimodal AI-based decision-support system for stroke classification during medicalemergency calls and compared its performance with the Emergency Medical Communication Centre (EMCC) operator baseline.
Methods:
We performed a retrospective diagnostic accuracy study using data from Bergen EMCC from 2019 and 2022. Two datasets were used: (1) 473 EMCC audio files with paired manual transcripts for fine-tuning a Norwegian Whisper-based automatic speech recognition (ASR) model (2019), and (2) 2,627 labelled emergency calls for classification model developmentand evaluation (2022). Structured clinical data included hospital data, laboratory measurements, and medications. Performance was assessed using area under the receiver operating characteristic curve (AUC), precision, recall, and F1-score. A streaming evaluation at 30, 60, and 90 ssimulated real-time use.
Results:
Fine-tuning improved ASR accuracy, reducing word error rate from 8.3% to 7.2%. The multimodal classification model achieved a stroke-class F1-score of 0.77, exceeding that of both the transcript-only model and the operator baseline (0.73). The macro-averaged F1-score was 0.87 and the AUC was 0.86. At an assumed 1.2% stroke prevalence, the model was estimated to generate approximately 26 false alerts per 1,000 calls, compared with 36 for the operator baseline. In the streaming evaluation, stroke F1-score increased from 0.68 at 30 s to 0.85 at 90 s, stabilising performance during the first minute.
Conclusions:
A multimodal AI-based decision-support system integrating emergency call transcripts with patient health recordsimproved stroke classification performance and may support EMCC operators during stroketriage. External validation in larger datasets is required before clinical implementation.