Related Experiment Video
Updated: Jul 6, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Applying large language models to spam detection in the Kazakh low-resource language setting
Kumisbek Mukhammed-Ali1, Shormakova Assem2
1Al-Farabi Kazakh National University, Almaty, 050040, Kazakhstan. mali.kumisbek@gmail.com.
Multilingual large language models (LLMs) show inconsistent zero-shot spam detection for Kazakh. However, parameter-efficient fine-tuning with LoRA significantly boosts accuracy, achieving over 0.99 for spam classification tasks.
Area of Science:
- Natural Language Processing
- Machine Learning
- Computational Linguistics
Background:
- Low-resource languages like Kazakh present challenges for NLP tasks due to data scarcity.
- Multilingual transformer models offer potential for cross-lingual transfer learning.
- Spam classification is a critical security task requiring robust language model performance.
Purpose of the Study:
- To evaluate the effectiveness of multilingual transformer-based large language models (LLMs) for spam classification in Kazakh.
- To assess the performance gains from supervised learning, particularly parameter-efficient fine-tuning using LoRA.
- To analyze the generalization capabilities of LLMs on low-resource languages for security applications.
Main Methods:
- Testing zero-shot and LoRA-based supervised learning on multilingual models (bert-base-multilingual-cased, distilbert-base-multilingual-cased, xlm-roberta).
- Evaluating performance on spam detection tasks for the Kazakh language.
- Analyzing misclassified instances to identify linguistic patterns influencing errors.
Main Results:
- Zero-shot accuracy of multilingual LLMs was inconsistent, particularly for the underrepresented spam class in Kazakh.
- Supervised learning with LoRA fine-tuning achieved high accuracy (above 0.99) across all tested models.
- XLM-R demonstrated superior performance with a macro-F1 score of 0.99, significantly reducing false negatives.
Conclusions:
- Limited annotated Kazakh data combined with LoRA fine-tuning effectively enhances multilingual LLM performance for spam detection.
- Multilingual models require supervised fine-tuning to overcome generalization issues in low-resource languages.
- Findings highlight the potential of efficient fine-tuning techniques for security-related NLP tasks in underrepresented languages.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
14:34A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026