Related Experiment Videos
Information extraction from full text scientific articles: where are the keywords?
Parantu K Shah1, Carolina Perez-Iratxeta, Peer Bork
1Biocomputing, European Molecular Biology Laboratory, Heidelberg, Germany. shah@embl.de
BMC Bioinformatics
|May 31, 2003
Summary
Scanning full text scientific articles for biological information is valuable. While abstracts have a high keyword density, other sections offer richer biological data.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Scientific Literature Analysis
Background:
- Current biological information extraction methods often focus solely on article abstracts.
- Full-text scientific articles represent a vastly underutilized data source.
- The relevance and utility of extracting data from full-text articles remain key questions.
Purpose of the Study:
- To evaluate the distribution and relevance of biological keywords across different sections of scientific articles.
- To determine if full-text scanning offers advantages over abstract-only analysis for biological data extraction.
Main Methods:
- Analysis of keyword distribution across standard scientific article sections: abstract, introduction, methods, results, and discussion.
- Comparative assessment of keyword density and biological relevance in each section.
Main Results:
- Significant heterogeneity in keyword content was observed across different sections of scientific articles.
- While abstracts exhibit the highest ratio of keywords to total words, they may not be the optimal source for all biological data.
Conclusions:
- Full-text scientific articles contain biologically relevant data beyond the abstract.
- Exploring sections like introduction, methods, results, and discussion is crucial for comprehensive biological information extraction.
- The effort invested in scanning full-text articles can yield significant biological insights.