Related Experiment Video
Updated: May 7, 2026

Measuring Statistical Learning Across Modalities and Domains in School-Aged Children Via an Online Platform and Neuroimaging Techniques
Published on: June 30, 2020
BABSA: A large scale bangla aspect based sentiment analysis dataset
1North South University (register @ northsouth edu), Plot # 15, Block # B, Bashundhara R/A, Dhaka 1229, Bangladesh.
Abstract:
Bangla Aspect-Based Sentiment Analysis (ABSA) research lacks large, high-quality, fine-grained resources despite the widespread use of Bangla in digital communication. To address this gap, we introduce BABSA, a 15,860-instance dataset designed to support aspect extraction and aspect-specific sentiment classification in Bangla. The dataset is compiled from five major sources: BanglaBook, SentNoB, EmoNoBa, Sazzed, and a web-scraped Bangla news corpus collected between January and June 2025. After initial preprocessing steps including deduplication, normalization, text cleaning, and formatting, all instances were manually annotated using a structured three-pass protocol. Annotation guidelines defined aspect term boundaries, multi-word aspect handling, implicit and explicit sentiment expressions, and domain-specific aspect conventions. Inter-annotator agreement was assessed using Cohen's kappa, with a final agreement score of 0.84 for both aspect spans and sentiment labels. BABSA spans 21 diverse domains such as product reviews, social commentary, politics, entertainment, and general discussion, providing broad linguistic and topical coverage. Each entry includes the original Bangla sentence, aspect terms with exact character-level spans, and sentiment labels (positive, neutral, or negative) associated with each aspect. Additional metadata such as source domain, sentence length, and aspect counts are included to support downstream analysis. The dataset is partitioned into stratified train-test splits and is accompanied by detailed documentation describing the collection strategy, annotation scheme, quality control procedures, and data format specifications. BABSA is intended for reuse in tasks such as aspect term extraction, aspect-sentiment labeling, multi-aspect analysis, cross-domain ABSA, and benchmarking Bangla language models. Its scale, domain diversity, and fine-grained structure make it suitable for training and evaluating both classical NLP models and modern large language models. The dataset and accompanying scripts are publicly available, enabling transparent reuse, reproducibility, and further extensions by the research community.
Related Concept Videos
Attitudes
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Mean Absolute Deviation
Let us consider a dataset containing the number of unsold cupcakes in five shops: 10, 15, 8, 7, and 10. Initially, calculate the sample mean. Then calculate the deviation, or the difference, between each data value and the mean. Next, the absolute values of these deviations are added and divided by the sample size to...
Sampling Distribution
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
