Related Experiment Video
Updated: Apr 18, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
QTLMiner: QTL database curation by mining tables in literature
Jing Peng1, Xinyi Shi2, Yiming Sun2
1College of Electronic and Information, Northeast Agricultural University, Harbin, China, School of Computer Science and Technology, Changchun University of Science and Technology, Changchun, China and The Key Lab of Soybean Molecular Design Breeding, Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences, Harbin, China College of Electronic and Information, Northeast Agricultural University, Harbin, China, School of Computer Science and Technology, Changchun University of Science and Technology, Changchun, China and The Key Lab of Soybean Molecular Design Breeding, Northeast Institute of Geography and Agroecology, Chinese Academy of Sciences, Harbin, China.
This study introduces a novel method for extracting quantitative trait locus (QTL) information from tables and text in biomedical literature, significantly reducing data curation efforts. The approach achieved high precision and recall in a soybean QTL database curation task.
Area of Science:
- Biomedical Informatics
- Genomics
- Data Science
Background:
- Biomedical literature contains extensive experimental results in figures and tables.
- Quantitative trait locus (QTL) information is often presented in tables within scientific papers.
- Existing text-mining methods primarily focus on unstructured text, neglecting valuable tabular data.
Purpose of the Study:
- To develop and evaluate a method for extracting QTL information from both tables and plain text in biomedical literature.
- To address the gap in current text-mining approaches by enabling the mining of structured data within tables.
- To reduce the labor-intensive nature of biological database curation.
Main Methods:
- A novel method was proposed to extract QTL information from heterogeneous and complex tables.
- Tables were converted into a structured database format.
- Information from plain text was integrated with the structured table data.
Main Results:
- The method was applied to curate a soybean QTL database.
- 2278 records were successfully extracted from 228 research papers.
- The method demonstrated a high precision rate of 96.9% and a recall rate of 83.3%, with an F-value of 89.6%.
Conclusions:
- The developed method effectively extracts QTL information from biomedical literature tables and text.
- This approach significantly enhances the efficiency and accuracy of biological database curation.
- The findings highlight the potential for automated knowledge extraction from structured data in scientific publications.
More Related Videos
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Quartile
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
Data Collection II
Chi-square Analysis
The chi-square test was developed by Pearson in 1990.
The first step of performing a Chi-square analysis is to establish a null hypothesis, which assumes that there is no real...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Data Collection by Survey
Data Collection I