Related Experiment Video
Updated: Feb 11, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Leveraging Wikipedia knowledge to classify multilingual biomedical documents
Marcos Antonio Mouriño García1, Roberto Pérez Rodríguez1, Luis Anido Rifón1
1Department of Telematics Engineering, University of Vigo, Campus Lagoas-Marcosende, 36310 Vigo, Spain.
This study introduces a novel method using Wikipedia for multilingual biomedical document classification. The approach, leveraging cross-language concept matching, outperforms existing methods, demonstrating Wikipedia
Area of Science:
- Biomedical Informatics
- Natural Language Processing
- Knowledge Representation
Background:
- Classifying biomedical documents across multiple languages is challenging.
- Existing methods often rely on machine translation or complex linguistic tools.
- Leveraging structured knowledge bases like Wikipedia can offer an alternative approach.
Purpose of the Study:
- To develop and evaluate a classifier for multilingual biomedical documents using Wikipedia knowledge.
- To introduce and assess the effectiveness of a cross-language concept matching technique.
- To compare the proposed method against machine translation and MetaMap-based classifiers.
Main Methods:
- Document representation using concept weights derived from Wikipedia.
- Cross-language concept matching via Wikipedia interlanguage links.
- Experimental validation on two custom multilingual biomedical corpora (ML-UVigoMED and EFSG-UVigoMED).
Main Results:
- The proposed Wikipedia-based classifier significantly outperforms machine translation and MetaMap-based approaches.
- The cross-language concept matching technique effectively enables multilingual classification.
- Superior performance was observed across benchmark experiments.
Conclusions:
- Wikipedia knowledge is highly advantageous for multilingual biomedical document classification.
- The proposed method offers a robust and effective solution for cross-lingual information retrieval in the biomedical domain.
- This approach provides a scalable and efficient alternative to traditional multilingual classification techniques.
More Related Videos
Related Concept Videos
Classifying Matter by Composition
According to its composition, the matter can be classified into two broad categories — pure substances and mixtures.
A pure substance is a form of matter that has a constant composition throughout with uniform properties. For example, any sample of sucrose has the same composition and same physical properties, such as melting point, color, and sweetness, regardless of the source from which it is isolated.
A mixture is composed of two or...
Classifying Matter by State
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.
Methods of Documentation V: CBE
In CBE, healthcare professionals establish predefined standards of practice that define what constitutes...
Formats for Nursing Documentation
Nursing Assessment Form:
• A nursing assessment form is a foundational document that captures detailed patient data from physical assessments and nursing histories.
• It includes patient demographics, medical history,...

