Related Experiment Video
Updated: May 22, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Layout-aware text extraction from full-text PDF of scientific articles
Cartic Ramakrishnan1, Abhishek Patnia, Eduard Hovy
1Information Sciences Institute, University of Southern California, 4676 Admiralty Way, Suite 1001, Marina del Rey, CA, 90292-6695, USA. cartic@isi.edu.
This study introduces a new system for extracting text from PDF scientific articles. The Layout-Aware PDF Text Extraction (LA-PDFText) system accurately identifies and classifies text blocks for improved text mining.
Area of Science:
- Computational Biology
- Bioinformatics
- Scientific Publishing
Background:
- Portable Document Format (PDF) is the standard for online scientific publications.
- Extracting text from PDFs in a layout-aware manner is challenging for text mining and biocuration systems.
- Existing methods struggle to accurately process the structural information within PDF scientific articles.
Purpose of the Study:
- To introduce the Layout-Aware PDF Text Extraction (LA-PDFText) system.
- To facilitate accurate, layout-aware text extraction from PDF research articles for text mining applications.
- To provide a baseline system for future research in multi-modal content extraction.
Main Methods:
- The LA-PDFText system employs a three-stage process: detecting contiguous text blocks using spatial layout, classifying blocks into rhetorical categories via a rule-based method, and stitching blocks in the correct order.
- Spatial layout processing is used to identify and group text blocks.
- A rule-based classifier categorizes text blocks into logical units such as title, abstract, and body text.
Main Results:
- The LA-PDFText system achieved high performance in identifying text blocks and classifying them into rhetorical categories, with Precision=0.96, Recall=0.89, and F1=0.91.
- Evaluated the accuracy of the block detection algorithm and compared text extraction accuracy against the PDF2Text system and PubMed Central.
- Demonstrated the system's effectiveness in extracting text from section-wise grouped blocks.
Conclusions:
- LA-PDFText is an open-source tool that enables accurate text extraction from full-text scientific articles.
- The system provides a foundational solution for enhancing biomedical text mining and biocuration.
- The source code is publicly available for further development and application.
Related Concept Videos
Extraction: Advanced Methods
Extraction: Effects of pH
Atomic Force Microscopy
The AFM Probe
The probe is regarded as the heart of any AFM setup and comprises the...
Scanning Electron Microscopy
Fundamental Principles
Accelerated...
Extraction: Partition and Distribution Coefficients
For extracting a solute from an aqueous phase into an organic...
Western Blotting
The technique begins with separating proteins from the sample using sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), followed by protein transfer, immunoblotting, and finally, protein detection.
