Related Experiment Video
Updated: Oct 9, 2025

04:25
Author Spotlight: Bridging Gaps in Anatomy and Establishing a Foundation for Algorithmic Studies
Published on: December 15, 2023
3.0K
Benchmarking table recognition performance on biomedical literature on neurological disorders
Tim Adams1, Marcin Namysl2,3, Alpha Tom Kodamullil1
1Fraunhofer Institute for Algorithms and Scientific Computing, Schloss Birlinghoven, Sankt Augustin 53757, Germany.
Bioinformatics (Oxford, England)
|December 22, 2021
Summary
A new benchmark dataset for complex biomedical tables was created to improve table recognition systems. This dataset aids in tuning and evaluating applications for extracting information from challenging scientific documents.
Area of Science:
- Biomedical informatics
- Natural Language Processing
- Data Extraction
Background:
- Table recognition systems extract quantitative data from documents.
- Existing systems struggle with complex biomedical tables due to limited benchmark data.
- Biomedical literature contains intricate tables requiring specialized extraction methods.
Purpose of the Study:
- To introduce a novel, curated benchmark dataset for evaluating table extraction in the biomedical domain.
- To address the scarcity of training and benchmark data for complex scientific tables.
- To facilitate the development and assessment of advanced table recognition applications.
Main Methods:
- Developed a highly curated benchmark dataset from a hand-curated literature corpus on neurological disorders.
- Evaluated state-of-the-art table extraction systems using the proposed benchmark.
- Introduced a new performance metric and improvements for evaluation procedures.
Main Results:
- The benchmark dataset enables tuning and evaluation of table extraction applications for complex biomedical tables.
- Evaluation of existing systems highlighted challenges in recognizing intricate table structures.
- The proposed evaluation metric and improvements enhance performance assessment.
Conclusions:
- The created benchmark dataset is crucial for advancing table recognition in the biomedical field.
- Openly accessible dataset and source code promote further research and development.
- Addressing challenges in complex table recognition is vital for scientific data extraction.

