Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

896
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
896
Classification of Illness01:17

Classification of Illness

8.6K
The meaning of illness is individualized to each person who experiences an alteration in health. In contrast, disease is a medical term indicating a pathological change in the structure and function of the body or mind. It is a condition that has specific symptoms and boundaries.
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
8.6K
Bias in Epidemiological Studies01:29

Bias in Epidemiological Studies

1.3K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:  
1.3K
Study Designs in Epidemiology01:20

Study Designs in Epidemiology

881
Epidemiological study designs are fundamental tools for investigating the distribution, determinants, and control of health conditions in populations. They help researchers understand the relationships between exposures and outcomes, and they broadly fall into two categories: "observational" and "experimental" studies.
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
881
Causality in Epidemiology01:21

Causality in Epidemiology

1.5K
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
1.5K
Confounding in Epidemiological Studies01:27

Confounding in Epidemiological Studies

573
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
573

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Large language models for cancer registry abstraction: a real-world evaluation across models, variables, and cancer types.

medRxiv : the preprint server for health sciences·2026
Same author

De-escalated adjuvant radiotherapy versus standard adjuvant treatment for human papillomavirus-associated oropharyngeal squamous cell carcinoma (MC1675): a phase 3, open-label, randomised controlled trial.

The Lancet. Oncology·2025
Same author

Recovering missing electronic health record mortality data with a machine learning-enhanced data linkage process.

Journal of the American Medical Informatics Association : JAMIA·2025
Same author

Iron Deficiency in Collegiate Athletes Obtaining Preparticipation Hemoglobinopathy Screening in the Upper Midwest.

Pediatric blood & cancer·2024
Same author

The Global Prevalence of Iron Deficiency in Collegiate Athletes: A Systematic Review and Meta-Analysis.

Pediatric blood & cancer·2024
Same author

Prognostic Factors in Limited-Stage Small Cell Lung Cancer: A Secondary Analysis of CALGB 30610-RTOG 0538.

JAMA network open·2024

Related Experiment Video

Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K

Epidemiologic Method Review at Scale: Assessing Charlson Comorbidity Versioning Using a Large Language Model.

Joshua T Fuchs1, Cara Johnson1, Nathan Foster1

  • 1University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.

Medrxiv : the Preprint Server for Health Sciences
|October 3, 2025
PubMed
Summary

A large language model identified Charlson Comorbidity Index (CCI) versions in over 31,000 studies. Findings reveal most studies incorrectly cite the original CCI, hindering accurate comorbidity assessment in modern research.

More Related Videos

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.5K
Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

2.3K

Related Experiment Videos

Last Updated: Jan 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.0K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.5K
Constructing and Visualizing Models using Mime-based Machine-learning Framework
06:19

Constructing and Visualizing Models using Mime-based Machine-learning Framework

Published on: July 22, 2025

2.3K

Area of Science:

  • Epidemiology
  • Medical Informatics
  • Computational Linguistics

Background:

  • The Charlson Comorbidity Index (CCI) is a standard metric in epidemiological research for assessing disease burden.
  • Numerous adaptations of the original CCI have emerged since 1987, creating ambiguity regarding current usage.
  • The specific versions of the CCI employed in research and their temporal trends are not well-documented.

Purpose of the Study:

  • To develop and validate a large language model (LLM)-based approach for automatically extracting Charlson Comorbidity Index (CCI) version information from scientific literature.
  • To analyze the utilization trends of different CCI versions in published research since 2012.
  • To address the ambiguity in CCI implementation due to the frequent citation of outdated versions.

Main Methods:

  • A large language model was trained and applied to a dataset of 31,767 research articles published from 2012 onwards.
  • The LLM was designed to detect and extract specific Charlson Comorbidity Index (CCI) version citations within the text of the articles.
  • The extracted data was analyzed to determine the frequency and trends of CCI version usage.

Main Results:

  • The analysis of 31,767 articles revealed that 63% of studies citing a single CCI method referenced only the original 1987 publication.
  • This reliance on the original CCI is problematic for contemporary research utilizing modern datasets.
  • The study demonstrates the feasibility and scalability of using LLMs for automated literature reviews of methodological tools.

Conclusions:

  • The widespread citation of the original Charlson Comorbidity Index (CCI) publication leads to significant ambiguity in its real-world application in current research.
  • An LLM-based methodology offers a scalable and efficient solution for reviewing the implementation of research methods over time.
  • This approach can enhance the accuracy and consistency of epidemiological studies by clarifying the specific tools being used.