Related Experiment Video
Updated: Jun 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the performance of multilingual models in answer extraction and question generation.
Antonio Moreno-Cediel1, Jesus-Angel Del-Hoyo-Gabaldon1, Eva Garcia-Lopez2
1Departamento de Ciencias de la Computación, Universidad de Alcalá, 28805, Alcalá de Henares, Spain.
This study enhances multiple-choice test generation in Spanish by fine-tuning Transformer models like mT5-base for Answer Extraction (AE) and Question Generation (QG). The mT5-base model, trained on a combined Spanish dataset, achieved superior performance, setting a benchmark for future research.
Area of Science:
- Natural Language Processing (NLP)
- Artificial Intelligence (AI)
- Machine Learning (ML)
Background:
- Multiple-choice test generation is a complex NLP task, particularly in non-English languages due to limited prior research.
- Transformer architectures have advanced Answer Extraction (AE) and Question Generation (QG) tasks.
- Existing research often lacks focus on Spanish language NLP challenges for automated test creation.
Purpose of the Study:
- To develop and evaluate improved models for Answer Extraction (AE) and Question Generation (QG) in Spanish.
- To investigate the efficacy of an answer-aware methodology for Spanish NLP tasks.
- To establish a performance benchmark for AE and QG models in Spanish using various evaluation metrics.
Main Methods:
- Fine-tuning three multilingual Transformer models: mT5-base, mT0-base, and BLOOMZ-560M.
- Utilizing three datasets: a Spanish translation of SQuAD, the SQAC dataset, and their union (SQuAD + SQAC).
- Evaluating model performance using metrics such as BLEU1-4, METEOR, ROUGE-L, CIDEr, SARI, GLEU, WER, and cosine similarity.
Main Results:
- The mT5-base model, fine-tuned on the combined SQuAD + SQAC dataset, demonstrated the best performance for AE and QG tasks.
- Models trained solely on the SQAC dataset also yielded competitive results, indicating dataset effectiveness.
- mT5-base outperformed similar research works based on standard evaluation metrics like BLEU, METEOR, and ROUGE-L.
Conclusions:
- The mT5-base model, when trained with an answer-aware methodology on a combined Spanish dataset, is highly effective for AE and QG.
- The study provides a valuable benchmark for future research in Spanish NLP for automated test generation.
- Further exploration with newer models and datasets is recommended to advance the field.
More Related Videos
Related Concept Videos
Improving Translational Accuracy
Translation
Translation Produces the Building Blocks of Life
Proteins are...
Language and Cognition
Extraction: Advanced Methods
Self-Evaluation: Self-Enhancement and Self-Verification
Complementation Tests
Organisms heterozygous for different mutations are crossed pairwise in all combinations. If present on different genes, the mutations can complement each other by providing the missing...

