Related Experiment Video
Updated: Sep 23, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
On cross-lingual retrieval with multilingual text encoders
Robert Litschko1, Ivan Vulić2, Simone Paolo Ponzetto1
1University of Mannheim, Mannheim, Baden-Wuerttemberg Germany.
State-of-the-art multilingual transformers struggle with unsupervised cross-lingual retrieval, often underperforming older methods. Supervised fine-tuning is necessary for sentence retrieval, and in-domain tuning improves document retrieval performance.
Area of Science:
- Natural Language Processing
- Information Retrieval
- Machine Learning
Background:
- Pretrained multilingual transformer models (e.g., mBERT, XLM) are the current standard for cross-lingual transfer.
- Cross-lingual word embedding spaces (CLWEs) were previously dominant but are now largely superseded.
- The effectiveness of these transformers for cross-lingual retrieval tasks remains an open question.
Purpose of the Study:
- To systematically evaluate state-of-the-art multilingual encoders for cross-lingual document and sentence retrieval.
- To compare their performance against earlier CLWE-based models in unsupervised settings.
- To investigate the impact of supervised fine-tuning and domain adaptation on retrieval performance.
Main Methods:
- Benchmarking unsupervised ad-hoc sentence- and document-level cross-lingual information retrieval (CLIR).
- Evaluating multilingual encoders fine-tuned on English relevance data for zero-shot transfer.
- Introducing and testing localized relevance matching for document retrieval.
- Conducting in-domain contrastive fine-tuning for language transfer.
Main Results:
- Multilingual encoders do not significantly outperform CLWEs for unsupervised document-level CLIR.
- State-of-the-art performance in sentence-level retrieval is achieved with supervised specialization, not off-the-shelf models.
- Supervised re-ranking rarely improves performance in zero-shot cross-lingual and domain transfer scenarios.
- In-domain contrastive fine-tuning is crucial for improving ranking quality.
Conclusions:
- Vanilla multilingual transformers are insufficient for unsupervised cross-lingual document retrieval.
- Supervised fine-tuning and domain adaptation are critical for achieving high performance in cross-lingual retrieval tasks.
- Significant differences exist between cross-lingual retrieval and monolingual retrieval transfer, indicating potential "monolingual overfitting".
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Encoding
Automatic processing involves the encoding of details like time, space, frequency, and the meaning of words, usually done without conscious...
Retrieval
Recall involves accessing information without cues, such as during an essay test, where individuals must retrieve facts and concepts from memory unaided. Another example is remembering the name of a colleague...
Language and Cognition
Improving Translational Accuracy
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Empathy