Related Experiment Video
Updated: Jan 17, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.2K
Detection of offensive content in the Kazakh language using machine learning and deep learning approaches
Milana Bolatbek1, Moldir Sagynay1, Shynar Mussiraliyeva1
1Department of Cybersecurity and Cryptology, Al-Farabi Kazakh National University, Almaty, Kazakhstan.
Peerj. Computer Science
|September 24, 2025
Summary
This study developed effective machine learning models to detect harmful content in the Kazakh language on social media. The research highlights the need for specialized natural language processing (NLP) tools for morphologically rich languages.
Area of Science:
- Computational Linguistics
- Artificial Intelligence
- Social Media Analysis
Background:
- Social media platforms face challenges in detecting destructive content like extremism and cyberbullying.
- Standard natural language processing (NLP) models are often inadequate for morphologically rich languages such as Kazakh.
- There is an urgent need for specialized tools to ensure online safety in the Kazakh digital space.
Purpose of the Study:
- To adapt and apply machine learning and deep learning techniques for detecting destructive content in the Kazakh language.
- To address the linguistic complexities of Kazakh, including its agglutinative structure and rich morphology.
- To improve the accuracy and effectiveness of content classification for online safety.
Main Methods:
- Utilized machine learning algorithms including logistic regression and support vector machines (SVM).
- Employed deep learning techniques such as long short-term memory (LSTM) networks.
- Combined n-gram features and stemming methods with machine learning for enhanced classification.
Main Results:
- Achieved high accuracy in classifying destructive content in the Kazakh language.
- Demonstrated the effectiveness of integrating n-gram and stemming with machine learning approaches.
- Validated the performance of adapted NLP models on Kazakh text data.
Conclusions:
- Developing language-specific NLP tools is crucial for handling the complexities of Kazakh.
- The study provides a successful framework for detecting harmful content in lesser-resourced languages.
- Findings contribute to enhancing online safety and combating digital extremism in Kazakh-speaking communities.