Related Experiment Video
Updated: Aug 30, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Improving neural machine translation with POS-tag features for low-resource language pairs
Zar Zar Hlaing1, Ye Kyaw Thu2,3, Thepchai Supnithi2
1Faculty of Information Technology, King Mongkut's Institute of Technology Ladkrabang, Bangkok, 10520, Thailand.
Integrating linguistic features into neural machine translation models significantly improves low-resource language translation. Shared-multi-source transformer models with these features achieved the best results for Thai and Myanmar language pairs.
Area of Science:
- Computational Linguistics
- Natural Language Processing
- Machine Translation
Background:
- Statistical machine translation (SMT) has benefited from linguistic features, but their application in neural machine translation (NMT) for low-resource languages remains underexplored.
- Low-resource languages like Thai and Myanmar present unique challenges for NMT systems due to limited data.
- Previous NMT models have not fully leveraged linguistic information for these specific language pairs.
Purpose of the Study:
- To propose and evaluate transformer-based NMT models incorporating linguistic features for Thai-Myanmar, Myanmar-English, and Thai-English translation.
- To investigate the impact of adding part-of-speech (POS) or universal part-of-speech (UPOS) tags to NMT models.
- To compare the performance of novel multi-source and shared-multi-source transformer models against baseline NMT architectures.
Main Methods:
- Development of transformer, multi-source transformer, and shared-multi-source transformer models.
- Integration of POS/UPOS tags into source, target, or both sides of the input data.
- Utilizing a standard transformer model and an Edit-Based Transformer with Repositioning (EDITOR) model as baselines for comparison.
Main Results:
- The incorporation of linguistic features demonstrably enhanced the performance of transformer-based NMT models for low-resource language pairs.
- Shared-multi-source transformer models augmented with linguistic features outperformed both baseline transformer and EDITOR models.
- Significant improvements were observed in Bilingual Evaluation Understudy (BLEU) and character n-gram F-score (chrF) metrics.
Conclusions:
- Linguistic features are crucial for improving NMT performance, especially in low-resource scenarios.
- The proposed shared-multi-source transformer architecture effectively utilizes linguistic information for enhanced translation quality.
- This research provides a viable approach for advancing machine translation capabilities for under-resourced languages.
Related Concept Videos
Improving Translational Accuracy
Tagging and Fusion Proteins
Translation
Translation is the process of synthesizing proteins from the genetic information carried by messenger RNA (mRNA). Following transcription, it constitutes the final step in the expression of genes. This process is carried out by ribosomes, complexes of protein and specialized RNA molecules. Ribosomes, transfer RNA (tRNA), and other proteins produce a chain of amino acids—the polypeptide—as the end product of translation.
Translation Produces the Building Blocks of...
Post-translational Translocation of Proteins to the RER
Targeting proteins to the ER
Hsp40 and Hsp70 chaperone molecules bind the translated proteins in the cytosol to prevent their folding. The chaperone binding helps to keep the signal...
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Initiation of Translation

