Related Experiment Video
Updated: Mar 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Zero-shot English-Assamese neural machine translation via pivot-based cross-lingual embedding alignment and transfer
1School of Computer Science Engineering and Technology, Bennett University, Greater Noida, Uttar Pradesh, 201310, India.
This study introduces a novel zero-shot translation framework for low-resource Assamese, utilizing linguistic similarities with Bengali. The approach significantly improves English-to-Assamese translation accuracy without direct parallel data.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Machine Translation
Background:
- Low-resource languages like Assamese present significant challenges for Neural Machine Translation (NMT) due to limited parallel corpora.
- Existing NMT models struggle with data scarcity, hindering effective translation development for underrepresented languages.
Purpose of the Study:
- To develop a novel zero-shot translation framework for English-to-Assamese translation.
- To leverage linguistic proximity between Assamese and Bengali to overcome data scarcity.
- To improve translation quality and efficiency for low-resource languages.
Main Methods:
- Utilized a multilingual transfer learning architecture (mBART) combined with cross-lingual embedding alignment (Procrustes optimization).
- Implemented subword tokenization, embedding projection, and dynamic vocabulary expansion to address out-of-vocabulary (OOV) challenges.
- Employed advanced language tagging mechanisms to reduce off-target translation errors.
Main Results:
- Achieved a BLEU score of 28.65 on an English-Assamese dataset, outperforming direct translation baselines by 7.53 points.
- Human evaluations demonstrated high adequacy (4.5/5) and fluency (4.6/5), with an 8.4% reduction in error rates.
- Demonstrated a 26.9% improvement in inference speed compared to baseline models.
Conclusions:
- Transfer learning and cross-lingual alignment are effective strategies for bridging resource gaps in NMT for linguistically proximate languages.
- The proposed framework successfully enables zero-shot translation for low-resource languages by exploiting shared semantic spaces.
- The methods developed significantly enhance translation quality, efficiency, and robustness for languages with scarce parallel data.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Associative Learning
Classical conditioning, also known...
Cross Product
The magnitude of the cross product is obtained by multiplying the magnitude of both the vectors and the sine of the angle between them. This means that a larger angle between the vectors will lead to a greater magnitude of the cross product.
Cotranslational Protein Translocation
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
Initiation of Translation
First, the initiator tRNA must be selected from the pool of elongator tRNAs by eukaryotic initiation factor 2 (eIF2). The initiator tRNA (Met-tRNAi) has conserved sequence elements including modified bases at...