Related Experiment Video
Updated: May 2, 2026

Mapping Alzheimer's Disease Variants to Their Target Genes Using Computational Analysis of Chromatin Configuration
Published on: January 9, 2020
Influence of feature encoding and machine learning algorithms on ancestry inference using autosomal STR profiles: A
Kohei Tomonari1, Yoshito Tomisaka2, Yoshio Yamaoka3
1Criminal Investigation Laboratory, Oita Prefectural Police, Oita, Oita, Japan; Department of Environmental and Preventive Medicine, Oita University Faculty of Medicine, Yufu, Oita, Japan.
Abstract:
Biogeographic ancestry inference provides valuable insights into forensic science for identifying unknown remains and analyzing trace evidence found at crime scenes, particularly when other useful information is lacking. In this study, we investigated the impact of two feature encoding strategies (raw allele values and one-hot encoding) and three machine learning algorithms (Random Forest (RF), XGBoost (XGB), and Support Vector Machine (SVM)) on population classification using simulated autosomal STR profiles. Simulations were conducted using 20 autosomal STR loci commonly included in commercial kits, focusing on East Asian (EA) and Southeast Asian (SEA) populations. XGB demonstrated strong performance with both raw allele values and one-hot encoding. In contrast, while SVM achieved comparable accuracy with one-hot encoding, its performance markedly declined when raw allele values were used. RF consistently yielded lower accuracy than the other two algorithms. We also investigated the limits of intra-continental population classification and found that accuracy generally plateaued at approximately 80-85 % for EA vs. SEA, Japanese vs. Chinese, and Japanese vs. Vietnamese classifications, with similar performance observed in a preliminary evaluation using real Japanese profiles. However, in the classification of the genetically close Japanese and Korean populations, a marked discrepancy was observed between the simulation results and those derived from the real profiles. These findings clarify how feature encoding interacts with machine learning in STR-based population classification and delineate practical limits for closely related Asian populations.
More Related Videos
11:49Enhanced Genetic Analysis of Single Human Bioparticles Recovered by Simplified Micromanipulation from Forensic ‘Touch DNA’ Evidence
Published on: March 9, 2015
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genetic Drift
Pedigree Analysis
Pedigree Analysis