Related Experiment Video
Updated: Jul 5, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Exploiting and integrating rich features for biological literature classification
Hongning Wang1, Minlie Huang, Shilin Ding
1State Key Laboratory of Intelligent Technology and Systems, Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China. whn03@mails.tsinghua.edu.cn
BMC Bioinformatics
|May 9, 2008
Summary
This study introduces a new feature selection method (TF*ML) for biological literature classification. Integrating diverse features significantly enhances classification accuracy, outperforming previous benchmarks.
Area of Science:
- Bioinformatics
- Computational Biology
- Text Mining
Background:
- Automated text classification is crucial for managing large-scale bioscience data.
- Biological literature contains numerous domain-specific features that can improve classification.
- Effective feature selection and integration are key challenges in biological literature classification.
Purpose of the Study:
- To develop an efficient feature selection and integration strategy for biological literature classification.
- To improve the performance of automated text classification in the bioscience domain.
Main Methods:
- Proposed a novel feature value schema: TF*ML.
- Utilized features ranging from domain-independent string features to domain-dependent semantic template features.
- Implemented proper integration strategies among different feature types.
Main Results:
- Achieved significant performance improvements in Area Under the Curve (AUC) by 11.5% and F-Score by 8.8% compared to previous methods.
- Outperformed the best results from the BioCreAtIvE 2006 benchmark.
- Demonstrated the effectiveness of the proposed TF*ML schema and feature integration.
Conclusions:
- Different feature types offer varying discriminative power in literature classification.
- Integrating domain-independent and domain-dependent features substantially boosts classification performance.
- Proper feature integration helps overcome overfitting issues related to data distribution.
Related Concept Videos
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
Methods of Classification and Identification
Bacterial identification relies on a diverse array of techniques to classify and understand microorganisms, each tailored to uncover specific characteristics. Traditional morphological approaches, while still valuable, are limited for closely related or structurally simple organisms. Modern methods integrate biochemical, serological, genetic, and advanced molecular tools to achieve greater accuracy.Morphological and Biochemical TechniquesMorphological characteristics, such as cell shape and...
Genome Annotation and Assembly
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
