Related Experiment Video
Updated: Oct 7, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
A multilingual annotated dataset for metadata extraction from Persian, Arabic and English scientific articles: header
Arash Yousefi Jordehi1, Reza Fallah1, Ebrahim Nikmard1
1Department of Computer Engineering, Faculty of Engineering, University of Guilan, Rasht, Iran.
Abstract:
This data article describes a multilingual annotated dataset for training and evaluating metadata extraction systems for scientific articles written in Persian, Arabic and English. The dataset consists of 635 JSON files, one per scientific article, holding the annotations of the article header and of the reference section. The article-level language distribution is 352 Persian (55.4 %), 142 English (22.4 %) and 141 Arabic (22.2 %). Of the 635 articles, 621 contain annotated reference sections; the remaining 14 contain header annotations only. In total, 17,832 individual reference strings are annotated, comprising 8476 Persian (47.5 %), 5741 Arabic (32.2 %) and 3615 English (20.3 %) strings. Header annotations cover seven metadata fields: title, authors, affiliations, abstract, keywords, venue information and DOI. Reference annotations cover twelve metadata fields: title, authors, journal or conference name, volume, issue number, page range, year, DOI, publisher, book identifier, organisation and reference type, alongside the raw reference string. The source articles were retrieved from open-access repositories covering >580 conferences and journals. Annotation was carried out over a period of more than twelve months by a team of ten annotators using a purpose-built Java labelling tool with color-coded tagging, and every annotated file was checked against its source PDF in a dedicated review phase. Each file is self-contained: it stores the annotated text spans verbatim together with their labels, so the original PDF is not required in order to use the data. The format can be consumed directly by sequence labelling methods such as conditional random fields, recurrent neural networks, transformer-based token classifiers and large language models. The data can be reused to develop and benchmark metadata extraction and citation parsing systems and to support indexing in digital libraries and scholarly search services covering Persian and Arabic literature.
Related Concept Videos
Tandem Mass Spectrometry
NMR Spectroscopy of Aromatic Compounds
Genome Annotation and Assembly
Mass Analyzers: Overview
Mass Spectrometry of Amines
Mass Spectrometry: Overview