Related Experiment Video
Updated: Aug 5, 2026

Application of DNA Barcoding to Identify Medicinal Plants
Published on: November 1, 2024
Harnessing Genomic Information for Identifying the Geographic Origin of Five North American Tree Species in Trade
Pauline Hessenauer1,2,3, Melanie Zacharias1,2,3, Julien Prunier4,5
1Département des Sciences du bois et de la Forêt, Faculté de Foresterie, Géographie et Géomatique Université Laval, Pavillon Abitibi-Price Québec Québec Canada.
Abstract:
Genomic tools for traceability of wood products offer a powerful tool to support sustainable forestry, fight illegal logging, and improve conservation efforts. Traditional identification methods (e.g., wood anatomy, spectroscopy) are limited in resolution or scope, but genomic approaches can infer both species identity and geographic origin. Using lodgepole pine (Pinus contorta) as a case study, we first compared three types of SNP datasets-random, adaptation-linked, and machine learning-selected-for their effectiveness to assign origin. While group-based assignment methods (e.g., geographic or genetic) performed well with a small number of groups, their accuracy declined sharply as the number of groups increased, though it remained above random expectations even with 281 groups. To address this limitation, we developed a coordinate-based methodological framework that directly predicts geographic origin from SNP marker sets of varying sizes. We applied this framework to four additional North American tree species (Populus trichocarpa, black cottonwood; P. tremuloides, quaking aspen; Picea mariana, black spruce; Pinus strobus, eastern white pine), which represent a range of evolutionary histories and population structures. Using both classical and machine learning algorithms, including Linear Model (LM), K-Nearest Neighbor (KNN), Random Forest (RF), and Gradient Boosting (GB), we predicted latitude and longitude with mean distance errors ranging from 15 to 383 km, depending on the species. Accuracy varied depending on species-specific attributes such as population structure and sampling density. KNN performed best in datasets with high sampling density, such as P. trichocarpa, while GB achieved superior performance in more genetically homogenous taxa like P. strobus. This flexible, data-driven approach enables precise traceability across tree species and supports forensic, regulatory, and certification uses. We outline practical guidelines emphasizing accurate taxonomic identification, broad sampling, and the use of ~2000-5000 random genomic markers. Although KNN and Gradient Boosting generally performed well, species-specific differences warrant testing multiple algorithms.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Modern Molecular Taxonomy

