Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches01:23

Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches

375
Biopharmaceutical studies constitute a vital field aiming to enhance drug delivery methods and refine therapeutic approaches, drawing upon diverse interdisciplinary knowledge. In research methodologies, the choice between controlled and non-controlled studies significantly influences the study's reliability and accuracy.
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
375
Relative Risk01:12

Relative Risk

1.7K
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
1.7K
Improving Translational Accuracy02:07

Improving Translational Accuracy

14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
3.5K
Actuarial Approach01:20

Actuarial Approach

274
The actuarial approach, a statistical method originally developed for life insurance risk assessment, is widely used to calculate survival rates in clinical and population studies. This method accounts for participants lost to follow-up or those who die from causes unrelated to the study, ensuring a more accurate representation of survival probabilities.
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
274
Hazard Rate01:11

Hazard Rate

375
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
375

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

The importance of gene polymorphism in familial inheritance of endometriosis.

International journal of gynaecology and obstetrics: the official organ of the International Federation of Gynaecology and Obstetrics·2026
Same author

<i>JAK2</i>V617F Mutation in Endothelial Cells of Patients with Atherosclerotic Carotid Disease

Turkish journal of haematology : official journal of Turkish Society of Haematology·2024
Same author

Different perspectives on translational genomics in personalized medicine

Journal of the Turkish German Gynecological Association·2022
Same author

Gene pathway analysis of the endometrium at the start of the window of implantation in women with unexplained infertility and unexplained recurrent pregnancy loss: is unexplained recurrent pregnancy loss a subset of unexplained infertility?

Human fertility (Cambridge, England)·2022
Same author

Correction to: Assessment of the Role of Nuclear ENDOG Gene and mtDNA Variations on Paternal Mitochondrial Elimination (PME) in Infertile Men: an Experimental Study.

Reproductive sciences (Thousand Oaks, Calif.)·2022
Same author

Assessment of the Role of Nuclear ENDOG Gene and mtDNA Variations on Paternal Mitochondrial Elimination (PME) in Infertile Men: An Experimental Study.

Reproductive sciences (Thousand Oaks, Calif.)·2022

Related Experiment Video

Updated: Jan 6, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.5K

Clinical Risk Computation by Large Language Models Using Validated Risk Scores.

Kaan Kara1, Tuba Gunel2

  • 1Department of Molecular Biology and Genetics, Institute of Graduate Studies in Sciences, Istanbul University, Istanbul, Türkiye.

Journal of Medical Systems
|September 30, 2025
PubMed
Summary

Large Language Models (LLMs) can reliably calculate clinical risk scores, enhancing healthcare workflows. GPT-4o-mini and Gemini 2.5 Flash showed high accuracy, though complex scores like Framingham remain challenging for AI.

Keywords:
Artificial IntelligenceClinical Decision SupportClinical Risk ScoresLarge Language ModelsMedical Informatics

More Related Videos

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.5K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

980

Related Experiment Videos

Last Updated: Jan 6, 2026

An R-Based Landscape Validation of a Competing Risk Model
05:37

An R-Based Landscape Validation of a Competing Risk Model

Published on: September 16, 2022

2.5K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.5K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

980

Area of Science:

  • Artificial Intelligence in Medicine
  • Clinical Informatics
  • Natural Language Processing

Background:

  • Large Language Models (LLMs) offer advanced natural language understanding for healthcare.
  • Direct LLM risk prediction is unreliable due to bias and data complexity.
  • Using LLMs to compute established clinical risk scores enhances validity and transparency.

Purpose of the Study:

  • To evaluate the accuracy of public LLMs in calculating validated clinical risk scores.
  • To compare the performance of GPT-4o-mini, DeepSeek v3, and Google Gemini 2.5 Flash.
  • To assess LLM reliability for enhancing clinical workflows through interpretable score calculation.

Main Methods:

  • Generated 100 diverse patient profiles as natural language clinical notes.
  • Utilized LLMs (GPT-4o-mini, DeepSeek v3, Gemini 2.5 Flash) to extract data and compute five clinical risk scores.
  • Compared LLM-computed scores against reference scores using accuracy, precision, recall, F1, and Pearson correlation.

Main Results:

  • GPT-4o-mini and Gemini 2.5 Flash demonstrated near-perfect agreement with reference scores for most clinical risk assessments.
  • DeepSeek v3 showed lower performance compared to GPT-4o-mini and Gemini 2.5 Flash.
  • All evaluated LLMs encountered difficulties with the complex Framingham Risk Score calculation.

Conclusions:

  • LLMs can accurately compute established clinical risk scores, offering a trustworthy alternative to direct AI risk prediction.
  • GPT-4o-mini and Gemini 2.5 Flash show promise for integrating into clinical workflows for risk score calculation.
  • Further development is needed to address LLM challenges with highly complex risk assessment formulas.