Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Authors' response to Tiffet et al.'s comment on "Performance and Reproducibility of Large Language Models in Named Entity Recognition: Considerations for the Use in Controlled Environments".

Drug safety·2025
Same author

Provision and Characterization of a Corpus for Pharmaceutical, Biomedical Named Entity Recognition for Pharmacovigilance: Evaluation of Language Registers and Training Data Sufficiency.

Drug safety·2023
Same author

Supervised Machine Learning-Based Decision Support for Signal Validation Classification.

Drug safety·2022
Same author

Atmospheric Correction Inter-comparison eXercise.

Remote sensing·2020
Same author

Adverse Events in Twitter-Development of a Benchmark Reference Dataset: Results from IMI WEB-RADR.

Drug safety·2020
Same author

Effects of salinity, temperature, and polarization on top of atmosphere and water leaving radiances for case 1 waters.

Applied optics·2012

Related Experiment Video

Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

502

Performance and Reproducibility of Large Language Models in Named Entity Recognition: Considerations for the Use in

Jürgen Dietrich1, André Hollstein2

  • 1Pharmaceuticals, Medical Affairs and Pharmacovigilance, Data Science and Insights, Bayer AG, Müllerstr. 178, 13353, Berlin, Germany. juergen.dietrich@bayer.com.

Drug Safety
|December 11, 2024
PubMed
Summary

Large language models (LLMs) show promise for healthcare, but reproducibility issues with GPT 3.5 and GPT 4 hinder their use in regulated systems. GPT models are recommended for generating training data proposals for improved performance.

More Related Videos

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K
Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
08:08

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese

Published on: April 1, 2016

9.3K

Related Experiment Videos

Last Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

502
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K
Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese
08:08

Comparing the Frequency Effect Between the Lexical Decision and Naming Tasks in Chinese

Published on: April 1, 2016

9.3K

Area of Science:

  • Artificial Intelligence in Healthcare
  • Natural Language Processing (NLP) for Medical Applications

Background:

  • Recent advancements in Artificial Intelligence (AI) have led to the development of large language models (LLMs) capable of human-like responses.
  • The potential application of LLMs in healthcare necessitates rigorous evaluation of their efficacy, reproducibility, and operability within controlled, regulated environments.

Purpose of the Study:

  • To assess the suitability of GPT 3.5 and GPT 4 for integration into GxP-validated systems.
  • To compare the performance of externally hosted GPT models against internally hosted LLMs.
  • To explore the zero-shot performance of LLMs for Named Entity Recognition (NER) and relation extraction, and to evaluate their potential for generating training data proposals.

Main Methods:

  • Reproducibility experiments were conducted to ascertain LLM viability in controlled settings.
  • Guided generation was employed to ensure consistent prompting across diverse models.
  • Few-shot learning and Quantized Low-Rank Adaptation (QLoRA) fine-tuning were utilized to enhance LLM performance.

Main Results:

  • Zero-shot GPT 4 performance was found to be comparable to a fine-tuned T5 model.
  • Zephyr demonstrated superior performance over zero-shot GPT 3.5, though fine-tuned T5 excelled in recognizing product combinations.
  • Both GPT variants exhibited a lack of reproducibility, a critical limitation for regulated environments.

Conclusions:

  • The reproducibility limitations of closed, proprietary LLMs like GPT variants may impede their use in regulated healthcare systems.
  • Despite reproducibility challenges, GPT models demonstrate strong Named Entity Recognition (NER) capabilities, making them valuable for generating annotation proposals to train other models.
  • Further research into internally hosted LLMs and fine-tuning techniques is warranted for reliable integration into validated systems.