Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Jul 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

Dictionary-Augmented Large Language Model Postprocessing for Bilingual Code-Switched Medical Speech Recognition:

Chanryeong Oh1, Yul Hwangbo1,2, Wonjoong Cheon3

  • 1Healthcare AI Team, National Cancer Center, Goyang-si, Gyeonggi-do, 10408, Republic of Korea, 821075293175.

Journal of Medical Internet Research
|July 8, 2026
PubMed
Summary

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Leaf-Specific Classification of Multi-Leaf Collimator Positioning Errors in Volumetric Modulated Arc Therapy Using a Convolutional Neural Network.

Journal of clinical medicine·2026
Same author

Temporal Trend of Cardiovascular Disease Burden Among Cancer Patients Between 2005 and 2022: Nationwide Population-Based Cohort Study in South Korea.

Korean circulation journal·2026
Same author

Shared Decision-Making for Determining Treatment Strategies in Low-Risk Thyroid Cancer: Protocol of a Multicenter Cluster-Randomized Trial (MAeSTro-SDM).

Journal of Korean medical science·2026
Same author

Bridging Photoacoustic and Protoacoustic Imaging: Material Heterogeneity Effects on Proton Range Verification Using Time-of-Flight Analysis.

Bioengineering (Basel, Switzerland)·2025
Same author

Log file-based quality assurance method for respiratory gating system.

Journal of applied clinical medical physics·2025
Same author

Enhancing Large Language Model Reliability: Minimizing Hallucinations with Dual Retrieval-Augmented Generation Based on the Latest Diabetes Guidelines.

Journal of personalized medicine·2024

A hybrid approach combining dictionary normalization and large language models (LLMs) significantly improved automatic speech recognition (ASR) accuracy for Korean-English medical speech. This method offers a practical solution for multilingual ASR challenges without retraining models.

Area of Science:

  • Medical Informatics
  • Natural Language Processing
  • Speech Recognition

Background:

  • Physician burnout is exacerbated by clinical documentation burdens, with significant time spent on electronic health records.
  • Automatic speech recognition (ASR) systems show promise for reducing documentation time but face challenges with Korean-English code-switching in medical settings.

Purpose of the Study:

  • To develop and evaluate a hybrid postprocessing strategy for ASR systems handling Korean-English code-switched medical speech.
  • The strategy combines medical terminology dictionary normalization with large language model (LLM)-based postprocessing to enhance ASR accuracy.

Main Methods:

  • A speech dataset was created from 23,652 nursing notes, comprising Korean, English, and numerals.
  • OpenAI's gpt-4o-transcribe model was used for initial speech recognition.
Keywords:
automatic speech recognitioncode-switchinglarge language modelmedical terminologynatural language processing

Related Experiment Videos

Last Updated: Jul 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

  • A hybrid postprocessing approach involved medical terminology dictionary normalization and evaluation of six LLMs (GPT and Claude variants) across different temperature settings.
  • Main Results:

    • The hybrid approach, particularly with Claude Sonnet 4, achieved a BERTScore of 0.9638 and a character error rate (CER) of 0.0820, a 64.9% reduction from baseline ASR (CER 0.2336).
    • Dictionary-based normalization consistently improved LLM-only postprocessing performance.
    • LLM-only postprocessing reduced CER by up to 36.09% compared to baseline.

    Conclusions:

    • A hybrid pipeline integrating dictionary normalization and LLM postprocessing significantly enhances ASR accuracy for Korean-English code-switched medical speech.
    • This modular framework provides a practical solution for multilingual ASR challenges without requiring model retraining.