Related Experiment Video
Updated: Jun 8, 2026

Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
A framework for semisupervised feature generation and its applications in biomedical literature mining
Yanpeng Li1, Xiaohua Hu, Hongfei Lin
1College of Computer Science and Technology, Dalian University of Technology, Dalian 116024, Liaoning, China, liyanpeng.lyp@gmail.com
This study introduces a feature coupling generalization (FCG) framework to create new features from unlabeled data. FCG enhances sparse features, significantly boosting performance in biomedical text mining tasks like named entity recognition and protein-protein interaction extraction.
Area of Science:
- Machine Learning
- Natural Language Processing
- Bioinformatics
Background:
- Feature representation is crucial for machine learning and text mining.
- Supervised learning methods often struggle with sparse features in labeled data.
Purpose of the Study:
- To develop a novel framework for generating informative features from unlabeled data.
- To enhance the utility of sparse features in supervised learning models.
- To improve performance in biomedical literature mining tasks.
Main Methods:
- Introduced the feature coupling generalization (FCG) framework.
- Identified example-distinguishing features (EDFs) and class-distinguishing features (CDFs).
- Generalized EDFs into higher-level features using their coupling with CDFs in unlabeled data (over 20 GB of PubMed abstracts).
Main Results:
- FCG effectively utilizes sparse features often ignored by traditional supervised learning.
- Significant performance improvements were observed: 7.8% in gene NER, 5.0% in PPIE, and 5.8% in GO annotation TC.
- Achieved high F-scores of 89.1 and 64.5, and a normalized utility of 60.1 on benchmark datasets.
Conclusions:
- The FCG framework successfully enriches sparse features by leveraging unlabeled data.
- This approach offers a powerful method for incorporating new information and boosting performance in various text mining applications, particularly in the biomedical domain.
More Related Videos
03:37Generating the Transcriptional Regulation View of Transcriptomic Features for Prediction Task and Dark Biomarker Detection on Small Datasets
Published on: March 1, 2024
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018