Related Experiment Video
Updated: Apr 21, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
15.8K
Adaptive semantic tag mining from heterogeneous clinical research texts
1Chunhua Weng, Ph.D., Associate Professor, Department of Biomedical Informatics, Columbia University, 622 W 168 Street, PH-20, New York, NY, 10032, USA,
Methods of Information in Medicine
|October 21, 2014
Summary
This study introduces an adaptive framework for mining frequent semantic tags (FSTs) from clinical research texts, improving recall and efficiency. The adaptable design shows consistent trends across diverse text types, enhancing clinical text mining capabilities.
Area of Science:
- Clinical Informatics
- Natural Language Processing
- Data Mining
Background:
- Heterogeneous clinical research texts contain valuable semantic information.
- Efficiently mining frequent semantic tags (FSTs) is crucial for knowledge discovery.
- Existing tag-mining algorithms may lack adaptability across different text types.
Purpose of the Study:
- To develop a flexible and adaptive framework for mining frequent semantic tags (FSTs) from diverse clinical research texts.
- To enhance the recall and efficiency of FST mining.
- To assess the framework's adaptability to various clinical text genres.
Main Methods:
- Developed a "plug-n-play" framework integrating unsupervised kernel algorithms with functional wrappers.
- Implemented temporal information identification and semantic equivalence detection as example wrappers.
- Compared performance against a baseline algorithm on ClinicalTrials.gov data and assessed adaptability on clinical data requests and trial protocols.
Main Results:
- Achieved a 12.8% increase in average recall and a 47.02% increase in speed compared to the baseline.
- Maintained high overlap (76.9%-100%) with baseline relevant FSTs across varying frequency thresholds.
- Observed consistent FST prevalence trends across ClinicalTrials.gov, data requests, and protocols, with saturation around 200 documents.
Conclusions:
- The proposed adaptive tag-mining framework is scalable and adaptable without compromising recall.
- The component-based architecture offers potential generalizability for other clinical text mining methods.
- This approach facilitates more robust and versatile analysis of clinical research data.

