Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Sample Size Calculation01:19

Sample Size Calculation

3.3K
Knowledge of the sample size is the first requirement to conduct random sampling or an experiment. The sample size is the total number of units, observations, or groups (in some cases) used to get the data to estimate a population parameter. As the name suggests, the sample size is that of the sample drawn from the population and differs from the population size.
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
3.3K
Improving Translational Accuracy02:07

Improving Translational Accuracy

2.6K
2.6K
Goodness-of-Fit Test01:16

Goodness-of-Fit Test

3.3K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
3.3K
Higher Mental Functions of the Brain: Language01:10

Higher Mental Functions of the Brain: Language

782
Language is a system of communication that allows the expression of thoughts, ideas, and feelings. The brain processes language in both hemispheres.
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
782
Detection of Gross Error: The Q Test01:00

Detection of Gross Error: The Q Test

6.0K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.0K
Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.5K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Improvement of persistent anuria by long-term percutaneous ventricular assist device management and successful bridge to left ventricular assist device implantation in a patient with acute myocardial infarction: a case report.

European heart journal. Case reports·2026
Same author

Performance of large language models and prompt engineering strategies for data extraction in systematic reviews.

Frontiers in digital health·2026
Same author

Use of Commercially Available Large Language Models to Generate Information Leaflets on Post-Intensive Care Syndrome: Clinical Utility Assessment.

JMIR formative research·2026
Same author

Single-cell RNA sequencing identifies M2-like macrophage polarization associated with mesenchymal stem cell treatment in a murine sepsis model.

Biomolecules & biomedicine·2026
Same author

Correction: Plasma apolipoprotein A-I is a causal protective factor in sepsis.

Scientific reports·2026
Same author

Claudin 4 Deletion Improves Gut Permeability and Survival in a Murine Model of Abdominal Sepsis.

Shock (Augusta, Ga.)·2026

Related Experiment Video

Updated: Jun 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

543

Performance of a Large Language Model in Screening Citations.

Takehiko Oami1, Yohei Okada2,3, Taka-Aki Nakada1

  • 1Department of Emergency and Critical Care Medicine, Chiba University Graduate School of Medicine, Chiba, Japan.

JAMA Network Open
|July 8, 2024
PubMed
Summary

Large language models (LLMs) show promise for citation screening in systematic reviews, demonstrating acceptable sensitivity and high specificity while significantly reducing screening time compared to conventional methods.

More Related Videos

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K
P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
06:09

P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation

Published on: September 8, 2023

557

Related Experiment Videos

Last Updated: Jun 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

543
A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K
P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation
06:09

P300-Based Brain-Computer Interface Speller Performance Estimation with Classifier-Based Latency Estimation

Published on: September 8, 2023

557

Area of Science:

  • Medical Informatics
  • Systematic Review Methodology
  • Artificial Intelligence in Healthcare

Background:

  • Systematic reviews are crucial for evidence-based medicine but are time-consuming.
  • Citation screening is a major bottleneck in the systematic review process.
  • Large language models (LLMs) offer potential for automating and improving efficiency in literature reviews.

Purpose of the Study:

  • To evaluate the accuracy and efficiency of a large language model (LLM) for title and abstract citation screening.
  • To compare LLM-assisted screening with the conventional manual screening method.

Main Methods:

  • A prospective diagnostic study evaluated GPT-4 Turbo for citation screening across 5 clinical questions for sepsis guidelines.
  • LLM screening decisions were based on predefined inclusion/exclusion criteria.
  • Performance metrics (sensitivity, specificity) and time efficiency were compared against manual screening.

Main Results:

  • LLM-assisted screening achieved a sensitivity of 0.75 and specificity of 0.99 in the primary analysis.
  • Post hoc prompt modifications improved sensitivity to 0.91 while maintaining high specificity (0.98).
  • LLM screening reduced processing time for 100 studies from 17.2 minutes to 1.3 minutes.

Conclusions:

  • LLM-assisted citation screening demonstrates acceptable sensitivity and high specificity.
  • This method significantly reduces the time required for literature screening in systematic reviews.
  • LLMs have the potential to enhance efficiency and decrease workload in systematic reviews.