Related Experiment Video
Updated: Feb 27, 2026

05:56
Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
3.3K
ANCHOLIK-NER: A benchmark dataset for Bangla regional named entity recognition
Bidyarthi Paul1, Faika Fairuj Preotee1, Shuvashis Sarker1
1Department of CSE, Southeast University, Dhaka, Bangladesh.
Plos One
|February 25, 2026
Summary
This study introduces ANCHOLIK-NER, the first dataset for Named Entity Recognition (NER) in Bangla regional dialects. Bangla BERT shows strong performance, highlighting the need for dialect-aware NLP models.
Area of Science:
- Computational Linguistics
- Natural Language Processing (NLP)
- Low-Resource Language Technologies
Background:
- Named Entity Recognition (NER) for regional Bangla dialects is underexplored.
- Existing NLP models struggle with unique linguistic features of dialects like Barishal, Chittagong, Mymensingh, Noakhali, and Sylhet.
- Lack of dedicated datasets and benchmarks hinders research in this area.
Purpose of the Study:
- Introduce ANCHOLIK-NER, the first benchmark dataset for NER in Bangla regional dialects.
- Provide baseline performance metrics for transformer-based models on this dataset.
- Facilitate future research in dialect-aware NLP for low-resource languages.
Main Methods:
- Created ANCHOLIK-NER dataset: 17,405 sentences, 101,817 words, 10 entity tags across 5 regions.
- Sourced data from public resources and used manual translations for entity alignment.
- Evaluated three transformer models: Bangla BERT, Bangla Bert Base, and BERT Base Multilingual Cased.
Main Results:
- Bangla BERT achieved the highest F1-scores across dialects: Mymensingh (82.27%), Barishal (81.48%), Sylhet (78.75%), Noakhali (78.50%), Chittagong (75.31%).
- Demonstrated strong performance in Mymensingh and Barishal dialects.
- Identified Chittagong dialect as more challenging due to significant variation.
Conclusions:
- ANCHOLIK-NER is a foundational resource for Bangla dialect NER.
- Bangla BERT offers promising performance but dialect-specific challenges remain.
- Future work should focus on dialect-aware adaptation and expanding dataset coverage.
Related Concept Videos
Regional Terms
16.5K
Regional terms describe anatomy by dividing the body parts into different regions that contain structures involved in contributing similar functions. Using these terms helps increase the accurate description and identification of the particular region of interest or region affected by the disease.
Primarily, the human body has two major regions, the axial and appendicular regions. The axial region comprises regions from the head to the abdomen and makes up the central body axis. In contrast,...
Primarily, the human body has two major regions, the axial and appendicular regions. The axial region comprises regions from the head to the abdomen and makes up the central body axis. In contrast,...
16.5K
Aggregates Classification
1.1K
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
1.1K
