Related Experiment Video
Updated: Sep 30, 2026

Technical Demonstration of Whole Genome Array Comparative Genomic Hybridization
Published on: August 5, 2008
Invited review: Theoretical concepts and practical applications of the APY algorithm for large-scale genomic
Mohammad Ali Nilforooshan1, Jorge Hidalgo2, Ivan Pocrnic3
1Livestock Improvement Corporation, Hamilton, New Zealand.
Abstract:
The increasing availability of genomic data has transformed modern genetic evaluation systems but has also introduced substantial computational challenges, particularly associated with the inversion of large genomic relationship matrices (G). In genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP), direct inversion of the G matrix becomes computationally infeasible beyond a threshold in the number of genotypes. Furthermore, G becomes numerically positive semidefinite or singular as the number of genotyped individuals approaches or exceeds the number of markers. The Algorithm for Proven and Young (APY) was developed to overcome this limitation and has since become a cornerstone of large-scale genomic evaluations. Alternative methods to APY exist, which do not require the inversion of G. Currently, applying these methods is relatively straightforward, but this was not the case in the past due to the lack of available software, high computational costs, and convergence issues. The structural similarity between ssGBLUP and BLUP made implementing APY easier than its counterparts. Those alternatives are presented in this review. This study provides a comprehensive synthesis of the theoretical foundations, methodological developments, and practical applications of APY. We present APY as a structured approximation of the inverse of a symmetric matrix with limited effective dimensionality, grounded in the biological properties of genomic data, including linkage disequilibrium, finite effective population size and independent chromosome segments. By conditioning the genomic information of a large set of non-core individuals on a smaller set of core individuals, APY yields a sparse approximate inverse of G, in which non-core individuals are conditionally independent given the core group. This formulation dramatically reduces computational and memory requirements while maintaining prediction accuracy when the core group adequately captures the dimensionality of genomic information. The review highlights that the most critical factors governing APY performance are the size and composition of the core group. Empirical and theoretical evidence shows that an optimal core size corresponds to the effective dimensionality of G, typically assessed via eigenvalue or singular value decomposition. Once this threshold is reached, further increases in core size offer negligible gains in accuracy. In single-breed populations, core composition becomes largely irrelevant beyond the optimal size, whereas in multi-breed and crossbred populations, careful core definition is required to ensure balanced representation and stable evaluations. Beyond its original role in GBLUP and ssGBLUP, APY is discussed as part of a broader class of approximate inverse and dimensionality-reduction methods, with conceptual links to eigenvalue truncation, reduced-rank models, and sparse Gaussian process approximations. The algorithm's applicability extends to alternative genomic models, preconditioning strategies for iterative solvers, and approximating other large symmetric matrices with limited effective rank. Finally, we place APY within the broader context of rapidly expanding genomic, phenomic, and multi-omic data sets. As biological data continue to grow in size and complexity, methods that exploit limited dimensionality and redundancy of information will remain essential. APY exemplifies how biologically informed approximations can deliver computationally efficient, robust, and interpretable solutions, and it provides a conceptual framework that is likely to influence future developments in quantitative genetics and large-scale data analysis.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genomics
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
