Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Genome-wide Association Studies-GWAS01:11

Genome-wide Association Studies-GWAS

15.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
15.1K
Single Nucleotide Polymorphisms-SNPs01:05

Single Nucleotide Polymorphisms-SNPs

17.7K
A single nucleotide polymorphism or SNP is a single nucleotide variation at a specific genomic position in a large population. It is the most prevalent type of sequence variation found in the human genome. Point mutations that occur in more than 1% of the population qualify as SNPs. These are present once every 1000 nucleotides on an average in the human genome. Replacement of a purine with another purine (A/G) or a pyrimidine with another pyrimidine (C/T) is known as a transition. In contrast,...
17.7K
Comparing Copy Number Variations and SNPs02:26

Comparing Copy Number Variations and SNPs

18.4K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
18.4K
Genetic Variation01:25

Genetic Variation

1.1K
Genetic variation is the diversity in DNA sequences found among individuals of the same species. This diversity is crucial for a species' survival because it helps organisms adapt to environmental changes. Genetic variation begins with fertilization, where an egg and sperm cell merge. Each of these cells carries 23 chromosomes, up to 46 in the fertilized egg. Chromosomes are long DNA strands that contain genes, the basic units of heredity.
Genes exist in different versions called alleles,...
1.1K
Genomics02:02

Genomics

39.3K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
39.3K
Behavioral Genetics and Its Designs01:23

Behavioral Genetics and Its Designs

850
Behavior genetics explores how genetic inheritance influences human behavior. It focuses on how genes, passed from parents to offspring, contribute to the development of behavioral traits and tendencies. This branch of genetics seeks to understand the complex interplay between inherited genetic factors and environmental influences in shaping our behaviors.
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
850

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Prevalence and recovery of taste dysfunction after stapedectomy in otosclerosis: a clinical study of 320 patients.

The Journal of laryngology and otology·2026
Same author

Interpretable machine learning for low-sample multi-omics: a case study of ferret vaccine response.

Bioinformatics advances·2026
Same author

The effect of multi-session cerebellar transcranial direct current stimulation on balance function in adults with chronic vestibular hypofunction.

American journal of otolaryngology·2026
Same author

Intravenous Hydroxocobalamin for Cyanide Poisoning From Smoke Inhalation: A Comprehensive Scoping Review.

Journal of the American College of Emergency Physicians open·2026
Same author

Comparison of single-cell sequencing technologies for allele-specific expression analysis in rabbit spermatids.

Genomics·2026
Same author

Democratising high performance computing for bioinformatics through serverless cloud computing: A case study on CRISPR-Cas9 guide RNA design with Crackling Cloud.

PLoS computational biology·2025

Related Experiment Video

Updated: Dec 12, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.6K

VariantSpark: Cloud-based machine learning for association study of complex phenotype and large-scale genomic data.

Arash Bayat1, Piotr Szul2, Aidan R O'Brien1

  • 1Health and Biosecurity, Commonwealth Scientific and Industrial Research Organisation (CSIRO), 11 Julius Ave North Ryde NSW 2113 Australia.

Gigascience
|August 8, 2020
PubMed
Summary

VariantSpark analyzes complex genetic traits by considering gene interactions, unlike traditional methods. This new framework efficiently processes large datasets to identify genetic variants associated with complex phenotypes.

More Related Videos

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
11:35

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA

Published on: August 21, 2016

13.3K
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

10.5K

Related Experiment Videos

Last Updated: Dec 12, 2025

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER
14:06

Detection of Rare Genomic Variants from Pooled Sequencing Using SPLINTER

Published on: June 23, 2012

15.6K
Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA
11:35

Screening for Functional Non-coding Genetic Variants Using Electrophoretic Mobility Shift Assay EMSA and DNA-affinity Precipitation Assay DAPA

Published on: August 21, 2016

13.3K
Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
05:53

Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry

Published on: June 21, 2018

10.5K

Area of Science:

  • Genetics
  • Computational Biology
  • Bioinformatics

Background:

  • Complex traits and diseases are often influenced by multiple genes (polygenic).
  • Polygenic risk scores (PRS) build on genome-wide association studies but do not account for gene-gene interactions (epistasis).
  • Analyzing epistasis in large datasets is computationally intensive, limiting current research.

Purpose of the Study:

  • To develop a scalable machine learning framework for association analysis of complex phenotypes.
  • To incorporate both additive genetic effects and epistatic interactions in risk models.
  • To enable analysis of population-scale genomic data with millions of variants.

Main Methods:

  • Developed VariantSpark, a distributed machine learning framework.
  • Implemented efficient multi-layer parallelization for whole-genome analysis.
  • Applied the framework to population-scale datasets (100,000 samples, 100,000,000 variants).

Main Results:

  • VariantSpark successfully performs association analysis for complex phenotypes involving polygenic and epistatic factors.
  • The framework scales efficiently to ultra-high-dimensional genomic data.
  • VariantSpark is 3.6 times faster than existing methods like ReForeSt.

Conclusions:

  • VariantSpark enhances the identification of genomic variants linked to complex phenotypes compared to traditional methods.
  • It is the first method capable of analyzing ultra-high-dimensional genomic data within a practical timeframe.
  • This advancement facilitates deeper understanding of the genetic architecture of complex diseases.