Related Experiment Video
Updated: Jul 18, 2026

08:53
Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
5.2K
Using Twitter to collect a multi-dialectal corpus of Albanian using advanced geotagging and dialect modeling
Ercan Canhasi1, Rexhep Shijaku1
1Faculty of Computer Science, University of Prizren, Prizren, Kosova.
Plos One
|November 27, 2023
Summary
This study created a geographically-informed Albanian National Corpus from Twitter data. Machine learning models accurately identified Albanian dialects, outperforming human annotators and revealing new linguistic patterns.
Area of Science:
- Computational Linguistics
- Sociolinguistics
- Natural Language Processing
Background:
- The need for a comprehensive, geographically-referenced Albanian National Corpus is significant for linguistic research.
- Existing resources lack sufficient dialectal and geographical representation, particularly from social media.
- Twitter data offers a rich, albeit challenging, source for capturing contemporary language use.
Purpose of the Study:
- To develop and categorize a geographically-informed, multi-dialectal Albanian National Corpus using Twitter data.
- To evaluate the efficacy of automated methods, including geotagging and machine learning, in corpus creation and dialect identification.
- To compare the performance of machine learning models against human annotators in dialect classification.
Main Methods:
- Automated data scraping and extraction from Twitter, identifying users with discernible locations and dialect usage.
- Application of advanced geotagging techniques for efficient corpus generation.
- Experimentation with diverse machine learning classification methodologies, including feature engineering and selection.
- Subjective assessment by human annotators for comparative analysis.
Main Results:
- Successfully assembled a publicly available, anonymized dataset of Albanian tweets with dialect annotations.
- Machine learning models demonstrated high proficiency in accurately differentiating Albanian dialects from individual tweets, surpassing human accuracy.
- Identified novel dialectal patterns previously unacknowledged in scientific literature.
Conclusions:
- Machine learning offers a powerful and accurate approach to dialect identification within the Albanian language using social media data.
- The developed corpus provides a valuable resource for future linguistic research on Albanian dialects.
- Automated methods significantly enhance the efficiency and accuracy of dialect corpus creation.
More Related Videos
Related Concept Videos
Extraction: Advanced Methods
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is formed in...
Impression Management Techniques IV: Altercasting
Altercasting is a strategic communication technique in which an individual imposes a specific identity or social role onto another person to influence their behavior and shape the interaction. By presuming a role—such as “responsible leader” or “patient person”—altercasting encourages the target to conform to that identity, often aligning their behavior with the expectations associated with the role. The power of this tactic lies in its subtlety; once a role is assigned, it becomes socially...

