Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Cancer Survival Analysis01:21

Cancer Survival Analysis

634
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
634
Improving Translational Accuracy02:07

Improving Translational Accuracy

14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy02:07

Improving Translational Accuracy

3.5K
3.5K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Association of <i>ABCB1</i> and <i>CYP2C19</i> polymorphisms with major adverse cardiovascular complications in patients taking clopidogrel.

Frontiers in cardiovascular medicine·2026
Same author

Cardiologists' perspectives on pharmacogenomics implementation in a hybrid health system: A qualitative study from the United Arab Emirates.

PloS one·2026
Same author

The diabetes exposome: interplay of environmental and genetic determinants in diabetes.

Human genomics·2026
Same author

AI-genomics synergy for drug repurposing in breast cancer: an interpretability-driven framework.

NPJ genomic medicine·2026
Same author

Correction: Combinational therapeutic strategies to overcome resistance to immune checkpoint inhibitors.

Frontiers in immunology·2026
Same author

Identification of a Novel VLDLR Variant in the First Report of CAMRQ1 From Africa: Expanding the Spectrum of Cerebellar Ataxia Syndromes.

Human mutation·2026

Related Experiment Video

Updated: Jan 11, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.3K

Leveraging AutoML to optimize dataset selection for improved breast cancer variants pathogenicity prediction.

Rahaf M Ahmad1, Noura AlDhaheri1, Mohd Saberi Mohamad1,2,3,4

  • 1Department of Genetics and Genomics, College of Medicine and Health Sciences, United Arab Emirates University, United Arab Emirates.

Computational and Structural Biotechnology Journal
|November 17, 2025
PubMed
Summary

Choosing the right dataset is crucial for accurately predicting breast cancer (BC) variant pathogenicity using automated machine learning (AutoML). A curated dataset combining cancer-specific and non-cancer data significantly improved prediction accuracy across multiple AutoML frameworks.

Keywords:
Automated machine learningBreast cancerDataset optimizationGenetic variantsPathogenicity prediction

More Related Videos

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.4K

Related Experiment Videos

Last Updated: Jan 11, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

7.3K
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.9K
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
07:41

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases

Published on: May 17, 2019

9.4K

Area of Science:

  • Genomic Medicine
  • Computational Biology
  • Machine Learning in Oncology

Background:

  • Breast cancer (BC) is a leading cause of cancer mortality globally, influenced by genetic and environmental factors.
  • Accurate prediction of genetic variant pathogenicity is vital for early detection, risk stratification, and personalized treatment of BC.
  • Current computational tools often lack disease specificity and generalizability for variant classification.

Purpose of the Study:

  • To benchmark the performance of different Automated Machine Learning (AutoML) frameworks (TPOT, H2O AutoML, MLJAR) for breast cancer variant pathogenicity prediction.
  • To evaluate the impact of dataset composition on the accuracy of pathogenicity prediction models.
  • To identify the optimal dataset for reliable BC-specific variant classification.

Main Methods:

  • Systematic benchmarking of four distinct variant datasets using three AutoML frameworks.
  • Evaluation of classification performance based on dataset composition.
  • Application of interpretability techniques (SHAP, permutation importance, LIME) to validate model transparency and biological relevance.

Main Results:

  • A curated dataset (Dataset-2) combining cancer-specific and non-cancer data consistently yielded the highest predictive performance across all tested AutoML frameworks.
  • H2O AutoML achieved a peak accuracy of 99.99% on the optimal dataset, with TPOT and MLJAR also demonstrating robust performance.
  • Feature importance analysis highlighted conservation scores and pathogenicity metrics as key predictors, with strong agreement across frameworks.

Conclusions:

  • Thoughtful dataset design, prioritizing disease-relevant and curated data, is critical for developing accurate machine learning models in genomic medicine.
  • The developed AutoML framework provides a scalable and interpretable approach for clinical prioritization of breast cancer variants.
  • This framework is adaptable for predicting pathogenicity in other genetic disorders, supporting precision diagnostics and personalized oncology.