Related Experiment Video
Updated: Sep 9, 2025

09:26
Quantification of Orofacial Phenotypes in Xenopus
Published on: November 6, 2014
9.8K
randPedPCA: rapid approximation of principal components from large pedigrees.
Hanbin Lee1, Rosalind Françoise Craddock2, Gregor Gorjanc2
1Department of Statistics, University of Michigan, Ann Arbor, MI, 48109, USA.
Genetics, Selection, Evolution : GSE
|August 28, 2025
Summary
Visualizing large pedigrees is now faster with the randPedPCA R package. It uses rapid pedigree principal component analysis (PCA) to efficiently analyze millions of records, overcoming computational challenges.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Modern breeding programs involve millions of pedigree records, posing visualization challenges.
- Pedigree data can be represented as relationship matrices, enabling visualization through principal component analysis (PCA).
- Naive PCA methods face computational and memory constraints with large pedigrees.
Purpose of the Study:
- To develop an efficient method for visualizing the structure of large pedigrees.
- To overcome computational limitations of traditional PCA for extensive pedigree datasets.
Main Methods:
- Developed the open-access R package randPedPCA for rapid pedigree PCA.
- Utilized sparse matrices and implicit matrix-vector multiplications with the inverse relationship factor.
- Implemented randomized singular value decomposition and Eigen decomposition via the RSpectra library.
Main Results:
- The randPedPCA package achieves speed-ups exceeding 10,000 times compared to naive PCA on simulated data.
- Randomized PCA methods enable analyses previously impossible due to computational constraints.
- Demonstrated analysis of a UK Kennel Club Labrador Retriever pedigree with nearly 1.5 million individuals, showing significant time reduction.
Conclusions:
- Leading principal components of pedigree matrices can be efficiently computed using randomized methods.
- Scatter plots of PCA scores provide intuitive visualizations for large pedigrees.
- This approach is substantially faster than traditional pedigree graph rendering for large datasets.
Related Concept Videos
Pedigree Analysis
85.1K
Overview
85.1K
Incomplete Dominance
25.4K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
25.4K
Hardy-Weinberg Principle
72.9K
Diploid organisms have two alleles of each gene, one from each parent, in their somatic cells. Therefore, each individual contributes two alleles to the gene pool of the population. The gene pool of a population is the sum of every allele of all genes within that population and has some degree of variation. Genetic variation is typically expressed as a relative frequency, which is the percentage of the total population that has a given allele, genotype or phenotype.
72.9K
Genetic Drift
40.6K
Natural selection—probably the most well-known evolutionary mechanism—increases the prevalence of traits that enhance survival and reproduction. However, evolution does not merely propagate favorable traits, nor does it always benefit populations.
40.6K
RACE - Rapid Amplification of cDNA Ends
6.5K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.5K

