Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

522
Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
522
Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

13.2K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
13.2K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

330
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
330

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Liver Tumor Segmentation with Deep Learning: A Comparative Analysis of CNN-, Transformer-, and YOLO-Based Models on the ATLAS MRI.

Diagnostics (Basel, Switzerland)·2026
Same author

A Comparative Analysis of the Mamba, Transformer, and CNN Architectures for Multi-Label Chest X-Ray Anomaly Detection in the NIH ChestX-Ray14 Dataset.

Diagnostics (Basel, Switzerland)·2025
Same author

CELM: An Ensemble Deep Learning Model for Early Cardiomegaly Diagnosis in Chest Radiography.

Diagnostics (Basel, Switzerland)·2025
Same author

A Fetal Well-Being Diagnostic Method Based on Cardiotocographic Morphological Pattern Utilizing Autoencoder and Recursive Feature Elimination.

Diagnostics (Basel, Switzerland)·2023

Related Experiment Video

Updated: Jun 14, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K

PyCaret for Predicting Type 2 Diabetes: A Phenotype- and Gender-Based Approach with the "Nurses' Health Study" and

Sebnem Gul1, Kubilay Ayturan1, Fırat Hardalaç1

  • 1Department of Electrical and Electronics Engineering, Faculty of Engineering, Graduate School of Natural and Applied Sciences, Gazi University, Ankara 06570, Turkey.

Journal of Personalized Medicine
|August 29, 2024
PubMed
Summary

Machine learning models can predict type 2 diabetes mellitus (T2DM) using phenotypic data. PyCaret successfully predicted T2DM, highlighting the importance of gender differences and family history in risk assessment.

Keywords:
PyCaretSHAP valuefeature importance plotmachine learningpredictiontype 2 diabetes mellitus

More Related Videos

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.0K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

661

Related Experiment Videos

Last Updated: Jun 14, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
12:18

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment

Published on: January 11, 2020

7.5K
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
07:31

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack

Published on: May 15, 2020

7.0K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

661

Area of Science:

  • Computational biology
  • Epidemiology
  • Machine learning applications in healthcare

Background:

  • Type 2 Diabetes Mellitus (T2DM) prediction is crucial for public health.
  • Machine learning (ML) techniques are increasingly used for disease prediction.
  • Phenotypic data offers a rich source for predictive modeling.

Purpose of the Study:

  • To evaluate the efficacy of PyCaret, an automated ML tool, in predicting T2DM.
  • To identify key phenotypic predictors of T2DM in male and female cohorts.
  • To assess the impact of gender on T2DM prediction models.

Main Methods:

  • Utilized PyCaret to apply 16 ML algorithms to phenotypic data from the "Nurses' Health Study" and "Health Professionals' Follow-up Study".
  • Analyzed separate male and female data subsets to identify best-performing models and influential features.
  • Evaluated model performance using AUC, accuracy, and precision metrics.

Main Results:

  • Ridge Classifier, Linear Discriminant Analysis, and Logistic Regression (LR) were optimal for males.
  • LR, Gradient Boosting Classifier, and CatBoost Classifier performed best for females.
  • Achieved AUCs of approximately 0.77 for males and 0.79 for females; accuracy and precision were around 0.70 and 0.71, respectively.
  • Key predictors included family history of diabetes, smoking status, and high blood pressure, with variations between genders.

Conclusions:

  • PyCaret effectively simplifies ML for T2DM prediction using phenotypic data.
  • Gender-specific analysis is essential for accurate T2DM risk prediction.
  • Future studies should consider integrating genotypic data with phenotypic data for enhanced early T2DM prediction.