Related Experiment Video
Updated: May 11, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Combining MEDLINE and publisher data to create parallel corpora for the automatic translation of biomedical text
Antonio Jimeno Yepes1, Elise Prieur-Gaston, Aurélie Névéol
1Lister Hill National Center for Biomedical Communications, US National Library of Medicine, Bethesda, USA. antonio.jimeno@gmail.com
Background:
Most of the institutional and research information in the biomedical domain is available in the form of English text. Even in countries where English is an official language, such as the United States, language can be a barrier for accessing biomedical information for non-native speakers. Recent progress in machine translation suggests that this technique could help make English texts accessible to speakers of other languages. However, the lack of adequate specialized corpora needed to train statistical models currently limits the quality of automatic translations in the biomedical domain.
Results:
We show how a large-sized parallel corpus can automatically be obtained for the biomedical domain, using the MEDLINE database. The corpus generated in this work comprises article titles obtained from MEDLINE and abstract text automatically retrieved from journal websites, which substantially extends the corpora used in previous work. After assessing the quality of the corpus for two language pairs (English/French and English/Spanish) we use the Moses package to train a statistical machine translation model that outperforms previous models for automatic translation of biomedical text.
Conclusions:
We have built translation data sets in the biomedical domain that can easily be extended to other languages available in MEDLINE. These sets can successfully be applied to train statistical machine translation models. While further progress should be made by incorporating out-of-domain corpora and domain-specific lexicons, we believe that this work improves the automatic translation of biomedical texts.
Related Concept Videos
Translation
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of Life
Complementary DNA
Improving Translational Accuracy
Improving Translational Accuracy
Transcription
Transcription is the process of synthesizing RNA from a DNA sequence by RNA polymerase. It is the first step in producing a protein from a gene sequence. Additionally, many other proteins and regulatory sequences are involved in the proper synthesis of messenger RNA (mRNA). Regulation of transcription is responsible for the differentiation of all the different types of cells and often for the proper cellular response to environmental signals.
Transcription Can Produce Different Kinds...
MicroRNAs

