Related Experiment Video
Updated: Jun 29, 2025

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Automatic categorization of self-acknowledged limitations in randomized controlled trial publications
Mengfei Lan1, Mandy Cheng2, Linh Hoang1
1School of Information Sciences, University of Illinois Urbana-Champaign, 501 Daniel Street, Champaign, 61820, IL, USA.
Researchers developed natural language processing (NLP) methods to automatically detect and categorize study limitations in randomized controlled trials (RCTs). This tool enhances scientific transparency by identifying reporting issues in publications.
Area of Science:
- Medical Informatics
- Clinical Trials Research
- Natural Language Processing
Background:
- Acknowledging study limitations is vital for scientific transparency and progress.
- Current reporting of limitations in scientific publications is often insufficient.
- Natural language processing (NLP) offers potential for automated checks to improve research transparency.
Purpose of the Study:
- To develop a dataset and NLP methods for detecting and categorizing self-acknowledged limitations in randomized controlled trial (RCT) publications.
- To improve the accuracy and efficiency of identifying and classifying study limitations.
- To enable large-scale analysis of limitation reporting in clinical research.
Main Methods:
- Created a data model with 15 categories and 24 sub-categories for limitation types.
- Annotated 1090 limitation instances across 200 full-text RCT publications.
- Fine-tuned BERT-based models (PubMedBERT) for sentence and type classification, incorporating data augmentation (EDA, PromDA).
- Applied the best-performing model to approximately 12,000 RCT publications.
Main Results:
- The PubMedBERT model for limitation sentence classification achieved an F1 score of 0.821.
- The best-performing limitation type classification model (PubMedBERT with PromDA) reached an F1 score of 0.7, a 2.7 percentage point improvement.
- Significant improvements in classification accuracy were observed with fine-tuning and data augmentation techniques.
Conclusions:
- Developed NLP models can support automated screening tools for journals to identify reporting issues.
- Automatic extraction of limitations can enhance peer review and evidence synthesis.
- This approach facilitates better searching and aggregation of evidence from clinical trial literature.
More Related Videos
12:55Multimodal Protocol for Assessing Metacognition and Self-Regulation in Adults with Learning Difficulties
Published on: September 27, 2020
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Related Concept Videos
Blinding
Self-Presentation: Self-Monitoring and Self-Handicapping
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Randomized Experiments
Simple randomization
Simple...
Blind Procedures
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...