Related Experiment Video
Updated: Jun 15, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Classifiers of Data Sharing Statements in Clinical Trial Records
Saber Jelodari Mamaghani1, Cosima Strantz1, Dennis Toddenroth1
1Medical Informatics, University Erlangen-Nuremberg, Germany.
Classifiers using language models can automatically identify available individual participant data (IPD) from clinical trial data-sharing statements (DSS). These models perform better when predicting manual labels than original categories, improving IPD discovery.
Area of Science:
- Computational linguistics
- Clinical trial data management
- Bioinformatics
Background:
- Digital individual participant data (IPD) from clinical trials is valuable for scientific reuse.
- Identifying available IPD requires interpreting textual data-sharing statements (DSS) in large databases.
- Computational linguistics, particularly pre-trained language models, offers potential for automating this interpretation.
Purpose of the Study:
- To evaluate the effectiveness of domain-specific pre-trained language models in classifying textual data-sharing statements (DSS).
- To compare classifier performance in reproducing original availability categories versus manually annotated labels for IPD.
- To assess the potential of these classifiers in aiding the automatic identification of available IPD.
Main Methods:
- Utilized a subset of 5,000 textual DSS from ClinicalTrials.gov.
- Developed and evaluated classifiers based on domain-specific pre-trained language models.
- Compared classifier performance against original availability categories and manually annotated labels using standard metrics.
Main Results:
- Classifiers trained to predict manually annotated labels outperformed those trained on original availability categories.
- This indicates that textual DSS contain information not captured by the original availability categories.
- The findings suggest significant potential for automated IPD identification.
Conclusions:
- Domain-specific pre-trained language models can effectively classify textual data-sharing statements.
- Manual annotations provide richer information for IPD availability classification than original categories.
- These advanced classifiers can significantly aid in the automatic discovery of reusable individual participant data from large clinical trial databases.
Related Concept Videos
Clinical Trials
There are four phases in a clinical trial. A phase one...
Clinical Trials: Overview
Hazard Ratio
For example, in a clinical trial...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Ethical Standards II
Nurses are entrusted with upholding various ethical principles and standards. Nurses forge solid therapeutic relationships using trust, empathy, autonomy, confidentiality, and professional competence.
Confidentiality is crucial, embodying respect for individual privacy...
Statistical Software for Data Analysis and Clinical Trials

