Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Improving Translational Accuracy02:07

Improving Translational Accuracy

9.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.4K
Accuracy, limits, and approximation01:28

Accuracy, limits, and approximation

445
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
445
Accuracy and Errors in Hypothesis Testing01:13

Accuracy and Errors in Hypothesis Testing

183
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
183
Reliability and Validity01:29

Reliability and Validity

12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
Stereotype Content Model02:16

Stereotype Content Model

14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Heuristics01:21

Heuristics

81
Heuristics are problem-solving strategies that use mental shortcuts to simplify decision-making. Unlike algorithms, which must be followed precisely to achieve a correct result, heuristics offer a general problem-solving framework. They save time and energy but can sometimes lead to less rational decisions.
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
81

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

OctoChemDB: An Aggregated Database for Small Molecule Identification Using High-Resolution MS Data.

Analytical chemistry·2026
Same author

Assessing the communication gap between AI models and healthcare professionals: Explainability, utility and trust in AI-driven clinical decision-making.

Artificial intelligence·2026
Same author

UpSMART: five years of digital innovation in cancer clinical research-achievements, challenges, and recommendations.

Frontiers in digital health·2025
Same author

Podocytes as novel sources of neutrophil serine proteases: expression and regulation by inflammatory molecular patterns.

Cellular and molecular life sciences : CMLS·2025
Same author

A Novel Role of Neutrophil Elastase in Podocyte Dysfunction Induced by High Glucose, PMA, and MDP.

Journal of cellular physiology·2025
Same author

Translating the machine; An assessment of clinician understanding of ophthalmological artificial intelligence outputs.

International journal of medical informatics·2025

Related Experiment Video

Updated: Jun 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

519

Large Language Models, scientific knowledge and factuality: A framework to streamline human expert evaluation.

Magdalena Wysocka1, Oskar Wysocki2, Maxime Delmas3

  • 1Digital Cancer Research, CRUK National Biomarker Centre, Manchester, United Kingdom; Department of Computer Science, University of Manchester, Manchester, United Kingdom.

Journal of Biomedical Informatics
|September 14, 2024
PubMed
Summary

Large Language Models (LLMs) show promise for scientific discovery but struggle with factual accuracy in biomedical knowledge tasks. Domain specialization and increased human feedback may improve their reliability as knowledge bases.

Keywords:
Antibiotic discoveryFactual knowledgeLarge language modelsRetrieval-augmented generation

More Related Videos

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.1K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K

Related Experiment Videos

Last Updated: Jun 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

519
Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.1K
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
05:48

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception

Published on: August 9, 2024

1.4K

Area of Science:

  • Biomedical Informatics
  • Artificial Intelligence in Science

Background:

  • Large Language Models (LLMs) offer potential for accelerating biomedical discovery by processing vast scientific literature.
  • Current LLM applications face challenges in accurately inferring and extracting complex biomedical information.

Purpose of the Study:

  • Introduce a novel framework to streamline the evaluation of factual scientific knowledge encoded by LLMs.
  • Assess the capabilities of eleven state-of-the-art LLMs in biomedical knowledge tasks, specifically antibiotic discovery.

Main Methods:

  • A three-step evaluation framework assessing fluency, alignment, coherence, factual knowledge, and specificity.
  • Task delegation between non-experts and domain experts to optimize evaluation efficiency.
  • Systematic assessment of LLMs on chemical compound definition and relation determination tasks.

Main Results:

  • LLMs demonstrate improved fluency but exhibit low factual accuracy and bias towards over-represented entities.
  • The reliability of LLMs as standalone biomedical knowledge bases is questionable.
  • The need for systematic evaluation frameworks for LLM-generated biomedical knowledge is highlighted.

Conclusions:

  • Current LLMs are not suitable for zero-shot biomedical factual knowledge retrieval.
  • Emerging properties suggest improved factuality with domain specialization, increased model scale, and human feedback.