Related Experiment Video
Updated: May 5, 2026

Mass Spectrometry-Guided Genome Mining as a Tool to Uncover Novel Natural Products
Published on: March 12, 2020
Mining Two Decades of Soybean Genomics Literature Using Rule-Based Text Mining: Chromosome-Resolved Patterns of Glyma
My Abdelmajid Kassem1, Dounya Knizia2, Khalid Meksem2
1Plant Genomics and Bioinformatics Lab, Department of Biological and Forensic Sciences, Fayetteville State University, Fayetteville, NC 28301, USA.
None:
Soybean (Glycine max [L.] Merr.) is a globally important crop with a rapidly expanding body of genomics literature driven by advances in sequencing and functional genomics. Thousands of studies reference soybean genes using standardized Glyma identifiers; however, systematic analyses of how these identifiers are distributed across chromosomes in the scientific literature remain limited. Here, we present a chromosome-resolved bibliometric analysis of soybean gene mentions using a reproducible rule-based text mining approach. PubMed abstracts published between December 2006 and December 2025 were mined for standardized Glyma gene identifiers using regular-expression-based entity extraction. A total of 377 PubMed records were retrieved, of which 340 abstracts (90.2%) contained at least one Glyma gene identifier. The median number of unique genes mentioned per abstract was 1, with a maximum of 14 genes reported in a single study. Our results reveal three major patterns. First, soybean genomics research remains predominantly gene-centric, with most abstracts referencing one or two genes. Second, apparent chromosome-level disparities exist in literature representation within the subset of studies using standardized Glyma identifiers, with chromosomes 3 and 16 exhibiting the highest frequencies of unique gene mentions. A Chi-square goodness-of-fit test confirmed that these differences deviate significantly from a uniform distribution (χ2 = 123.71, p < 0.001), indicating non-random patterns of gene reporting. Third, a small subset of genes dominates the literature, while the majority of annotated genes are mentioned infrequently, reflecting a long-tailed distribution of research attention. This analysis captures reporting patterns in studies that explicitly use standardized Glyma identifiers and therefore represents a defined subset of the broader soybean genomics literature. Within this scope, the findings highlight uneven adoption of standardized gene nomenclature and chromosome-level differences in research emphasis. More broadly, this study demonstrates the utility of transparent, rule-based text mining approaches for large-scale bibliometric analyses in plant science and provides a scalable framework for comparative analyses across crop species.
More Related Videos
07:34Author Spotlight: Soybean Hairy Root Transformation for the Analysis of Gene Function
Published on: May 5, 2023
10:40Comprehensive Workflow for the Genome-wide Identification and Expression Meta-analysis of the ATL E3 Ubiquitin Ligase Gene Family in Grapevine
Published on: December 22, 2017
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Genome Annotation and Assembly
Genetic Lingo
Cis-regulatory Sequences