Related Experiment Video
Updated: Jan 12, 2026

Author Spotlight: Investigating the Role of Repetitive DNA Misregulation in Cancer Initiation and Immunotherapy Resistance
Published on: December 13, 2024
Optimizing the performance of large genomic evaluations through data truncation in Angus cattle
Zuleica Trujano1,2, Andre Garcia2, Kelli Retallick2
1Department of Animal and Dairy Science, University of Georgia, Athens, GA 30602.
Abstract:
Single-step GBLUP provides accurate genomic breeding values (GEBV) for populations of any size. However, in large genomic models, computing time is costly, raising the question of whether using the full dataset is truly beneficial or if fewer data can achieve similar results while reducing computational costs. In this study, we aimed to assess the impact of data truncation on computing time and prediction accuracy in the American Angus growth model. The traits analyzed were birth weight (BW), weaning weight (WW), and post-weaning gain (PWG). The initial dataset included 12,802,165 phenotyped animals, 1,570,859 genotyped animals, and a total of 15,082,643 individuals in the pedigree. In phenotypic data truncation, we removed phenotypes for animals born before 1985 (P-1985), 1995 (P-1995), 2005 (P-2005), or 2015 (P-2015), with P-all retaining all records. In the genotypic data truncation, we excluded genotyped animals without records or progeny without records (G-info), whereas G-all included all genotyped animals. Predictions for genotyped animals excluded from the main evaluation were obtained as indirect predictions (IP). We validated the GEBV using the LR and predictive ability methods. The LR prediction accuracy across the scenarios was in the range of 0.62-0.63 (BW), 0.74-0.77 (WW), and 0.72-0.74 (PWG). Predictivity for P-all, P-1985, P-1995, and P-2005 was 0.51 for BW, 0.47 for WW, and 0.35 for PWG. The values for P-2015 were 0.01 lower than these. Correlations GEBV-IP were ≥ 0.99 for P-1985, P-1995, and P-2005. GEBV and IP had similar means and accuracies. The results showed that moderate phenotypic and genotypic data truncation (P-2005/G-info) was suitable, reducing computing time by 66% without compromising prediction accuracy and the model's ability to predict future phenotypes. This outcome reflected the limited influence of old data on the predictions of young animals and the minimal contribution of genotyped young animals without own or progeny records to the predictions of their relatives. Indirect predictions provided a fast, reliable way to predict genetic merit for non-informative animals excluded in the G-info scenario. Data truncation can preserve prediction accuracy with no impact on dispersion, particularly when phenotypic and genotypic datasets are large (robust data structure), genotyping is non-selective, traits have medium to high heritability, and pedigree depth is restricted to three or four generations.
More Related Videos
05:53Candidate Gene Testing in Clinical Cohort Studies with Multiplexed Genotyping and Mass Spectrometry
Published on: June 21, 2018
09:32An Array-based Comparative Genomic Hybridization Platform for Efficient Detection of Copy Number Variations in Fast Neutron-induced Medicago truncatula Mutants
Published on: November 8, 2017
Related Concept Videos
Truncation in Survival Analysis
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Pedigree Analysis
Incomplete Dominance
Evolutionary Relationships through Genome Comparisons
Genome Size and the Evolution of New Genes