Related Experiment Video
Updated: Apr 18, 2026

A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Microtask crowdsourcing for disease mention annotation in PubMed abstracts
Benjamin M Good1, Max Nanis, Chunlei Wu
1Molecular and Experimental Medicine, The Scripps Research Institute, 10550 N. Torrey Pines Rd., La Jolla, CA, 92037, USA. bgood@scripps.edu.
Amazon Mechanical Turk (AMT) effectively annotates biomedical text for disease mentions, achieving high F-measures. This crowdsourcing approach accelerates corpus creation for biological natural language processing (BioNLP) research.
Area of Science:
- Biomedical Natural Language Processing (BioNLP)
- Computational Biology
- Data Annotation
Background:
- Developing large, annotated corpora is crucial for advancing BioNLP and machine learning in biomedical text analysis.
- Traditional expert annotation is time-consuming and costly.
- Microtask crowdsourcing platforms show potential for generating high-quality annotations.
Purpose of the Study:
- To investigate the efficacy of Amazon Mechanical Turk (AMT) for annotating disease mentions in PubMed abstracts.
- To refine a crowdsourcing protocol for accurate biomedical text annotation.
- To benchmark the protocol against the NCBI Disease corpus.
Main Methods:
- Utilized AMT workers to annotate disease mentions in 593 PubMed abstracts.
- Developed and iterated on a crowdsourcing protocol.
- Merged annotations using a simple voting method.
- Benchmarked results against the NCBI Disease corpus gold standard.
Main Results:
- Achieved an overall F-measure of 0.872 (precision 0.862, recall 0.883) compared to the gold standard.
- Annotation quality, measured by F-measure, improved with more workers per task, with diminishing returns beyond 8 workers.
- Successfully annotated 593 documents in 9 days with 145 workers at a low cost per abstract.
Conclusions:
- Microtask crowdsourcing via AMT is a viable and efficient method for generating well-annotated biomedical corpora.
- This approach can significantly accelerate the creation of datasets for BioNLP research.
- The developed protocol demonstrates the potential of crowdsourcing for large-scale biomedical text annotation.
More Related Videos
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025