Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

533
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
533
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

1.6K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
1.6K
Hazard Ratio01:12

Hazard Ratio

705
The hazard ratio (HR) is a widely used measure in clinical trials to compare the risk of events, such as death or disease recurrence, between two groups over time. It reflects the ratio of hazard rates—the instantaneous risk of the event occurring—between a treatment group and a control group. This measure provides valuable insights into the relative effectiveness of a treatment by assessing how the risk of an event differs between the two groups.
For example, in a clinical trial...
705
Randomized Experiments01:13

Randomized Experiments

9.3K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
9.3K
Study Designs in Epidemiology01:20

Study Designs in Epidemiology

1.5K
Epidemiological study designs are fundamental tools for investigating the distribution, determinants, and control of health conditions in populations. They help researchers understand the relationships between exposures and outcomes, and they broadly fall into two categories: "observational" and "experimental" studies.
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
1.5K
Strategies for Assessing and Addressing Confounding01:25

Strategies for Assessing and Addressing Confounding

532
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
532

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Regional variations in surgeries for carpal tunnel syndrome and ulnar nerve disorders: A registry-based study in Finland.

Scandinavian journal of surgery : SJS : official organ for the Finnish Surgical Society and the Scandinavian Surgical Society·2026
Same author

An Acta Orthopaedica educational article: Choosing treatment with an elderly patient having a distal radius fracture.

Acta orthopaedica·2026
Same author

Interventions to improve neonatal and infant intubation success: a meta-analysis.

European journal of pediatrics·2026
Same author

Effect of Intraoperative Regional Anesthesia on Postoperative Outcomes in Pediatric Cardiac Surgery-A Systematic Review of Randomized Controlled Trials.

Paediatric anaesthesia·2026
Same author

Perioperative use of disease-modifying anti-rheumatic drugs (DMARDs) in people with inflammatory arthritis.

The Cochrane database of systematic reviews·2026
Same author

Surfactant-Budesonide Combination to Prevent Death or Bronchopulmonary Dysplasia: A Systematic Review and Meta-Analysis.

Neonatology·2026

Related Experiment Video

Updated: Mar 31, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.7K

Large language models for risk-of-bias assessment in randomised clinical trials-a comparative validation study.

Lauri Nyrhi1, Ville Ponkilainen1, Juho Laaksonen2

  • 1Department of Orthopaedics and Traumatology, Tampere University Hospital, Finland; Faculty of Medicine and Health Technology, Tampere University, Finland.

Ebiomedicine
|March 29, 2026
PubMed
Summary

Large language models (LLMs) show limited reliability for autonomous risk of bias assessment in randomized controlled trials. Human supervision is crucial for current LLM applications in evidence synthesis.

Keywords:
Artificial intelligenceLarge language modelMethodologyRisk of bias

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

8.3K

Related Experiment Videos

Last Updated: Mar 31, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.7K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

8.3K

Area of Science:

  • Artificial Intelligence
  • Medical Informatics
  • Clinical Trials

Background:

  • Large language models (LLMs) are increasingly used for evidence synthesis.
  • Risk of bias (RoB) assessment is crucial but time-consuming in trial evaluation.
  • Existing LLM studies show variable reliability for RoB screening.

Purpose of the Study:

  • To evaluate the accuracy and consistency of reasoning-enabled LLMs for RoB screening in randomized trials.
  • To compare the performance of four LLMs against human judgments.
  • To determine the potential of LLMs to reduce reviewer workload in evidence synthesis.

Main Methods:

  • A comparative validation study assessed four LLMs (ChatGPT o3, DeepSeek v3, Google Gemini Flash 2.0, Grok 3).
  • LLMs were prompted with full-text randomized clinical trial articles and protocols from two corpora (RoB 1 and RoB 2).
  • Interobserver reliability (Cohen κ) and diagnostic accuracy metrics were primary and secondary outcomes, compared to published human RoB judgments.

Main Results:

  • Interobserver agreement varied across LLMs, with DeepSeek v3 showing higher agreement (κ 0.39) on RoB 1.
  • Agreement was lower on RoB 2, with ChatGPT o3 (κ 0.06) and Gemini (κ 0.13) showing the lowest.
  • Diagnostic performance was limited, with sensitivity ranging from 0.05-0.55 and models consistently over-flagging concerns.

Conclusions:

  • Current LLMs are not reliable for fully autonomous RoB assessment.
  • DeepSeek v3 and ChatGPT o3 showed the best performance on RoB 1, but RoB 2 performance was modest.
  • LLMs require supervision and further improvements for safe stand-alone deployment, with potential use for triage or as a second assessor.