Related Experiment Video
Updated: Sep 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A survey on text classification: Practical perspectives on the Italian language.
Andrea Gasparetto1, Alessandro Zangari1, Matteo Marcuzzo1
1Department of Management, Ca' Foscari University, Venice, Italy.
Deep learning for text classification faces challenges in non-English languages like Italian due to data scarcity and computational costs. This study explores these issues and proposes solutions for linguistically inclusive natural language processing.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Machine Learning
Background:
- Deep learning has revolutionized text classification, primarily in English.
- Limited resources and linguistic complexities hinder adoption in other languages.
- Italian language text classification lacks extensive research and benchmarked datasets.
Purpose of the Study:
- To survey challenges in applying modern text classification to non-English languages, focusing on Italian.
- To highlight issues of dataset scarcity and computational expense.
- To propose a linguistically inclusive approach to text classification.
Main Methods:
- Comparative analysis of dataset availability for Italian and French.
- Application of representative text classification methods to custom multilabel datasets.
- Datasets created in Italian, French, and English for practical scenario simulation.
Main Results:
- Identified significant challenges in Italian text classification, including data scarcity.
- Demonstrated the impact of linguistic variations on model performance.
- Compared computational costs of modern approaches across languages.
Conclusions:
- Addressing data scarcity and computational challenges is crucial for non-English text classification.
- Further research is needed for linguistically inclusive and equitable NLP development.
- Future work should focus on creating diverse datasets and efficient models for under-resourced languages.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
06:48Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Types of Surveys
Classification of Systems-II
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Language and Cognition