Related Experiment Video
Updated: Mar 27, 2026

08:27
Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
7.3K
BanglaEcomReviewCorpus: A dataset for e-commerce product review sentiment analysis
Umme Ayman1, Md Tanvir Ahmed Akash1, Taslima Akhter1
1Department of Computer Science and Engineering. Daffodil International University, Bangladesh.
Data in Brief
|March 26, 2026
Summary
A new dataset of 8685 online customer reviews from Bangladeshi e-commerce sites enables advanced sentiment analysis. This resource supports improved business strategies and AI training through natural language processing (NLP) insights.
Area of Science:
- Computational Linguistics
- Data Science
- E-commerce Analytics
Background:
- Customer feedback is vital for e-commerce success, influencing business strategies, service quality, and product innovation.
- Understanding consumer behavior and preferences is essential for online retailers.
- Analyzing diverse customer feedback requires comprehensive and well-structured datasets.
Purpose of the Study:
- To introduce a novel, publicly available dataset of 8685 labeled customer feedback items from various Bangladeshi e-commerce platforms.
- To facilitate advanced sentiment analysis and the exploration of consumer behavior in the e-commerce domain.
- To support interdisciplinary research in areas such as sociology, linguistics, and psychology.
Main Methods:
- Data collection from popular e-commerce websites (Daraz, Bikroy.com, Picabbo, Shajgoj, etc.).
- Labeling of 8685 feedback items into positive, negative, and neutral sentiment categories.
- Application of statistical analyses (summary statistics, histograms) and linguistic pattern extraction (unigrams, bigrams, trigrams).
- Utilizing visualizations like word clouds for structural and linguistic diversity insights.
- Ensuring data integrity through rigorous collection, anonymization, and preprocessing techniques.
Main Results:
- A balanced dataset with 3012 positive, 2881 negative, and 2792 neutral feedback entries.
- Identification of linguistic patterns and trends within customer reviews.
- Demonstration of the dataset's utility for natural language processing (NLP) tasks.
- Provision of insights into consumer behavior and preferences across different e-commerce platforms.
Conclusions:
- The dataset is a valuable resource for advancing sentiment analysis and improving business strategies in e-commerce.
- It enables deeper insights into customer behavior, aiding product and service development.
- The dataset supports academic teaching, artificial intelligence (AI) training, and collaborative research across multiple disciplines.
Related Concept Videos
Review and Preview
12.1K
Data are individual items of information obtained from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete. Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random...
12.1K
Review and Preview
8.9K
In statistics, several tools are used to interpret the data. Measures of central tendency represent the characteristics of the data, such as mean, median, and mode. Additionally, measures of variance like standard deviation and range are used to find the spread of data from the mean. Relative standing measures the distance between data locations. Commonly used measures of relative standings are percentile, z score, and quartiles.
Percentiles are a type of fractile that partition data into...
Percentiles are a type of fractile that partition data into...
8.9K
Regression Analysis
8.9K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
8.9K
Goodness-of-Fit Test
9.4K
The goodness-of-fit test is a type of hypothesis test which determines whether the data "fits" a particular distribution. For example, one may suspect that some anonymous data may fit a binomial distribution. A chi-square test (meaning the distribution for the hypothesis test is chi-square) can be used to determine if there is a fit. The null and alternative hypotheses may be written in sentences or stated as equations or inequalities. The test statistic for a goodness-of-fit test is given as...
9.4K
Stereotype Content Model
15.7K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.7K
Data Collection by Survey
9.5K
The systematic method of obtaining and analyzing accurate information of a population is called data collection. A survey is a standard method of data collection that involves collecting information from a target human population about their experience, opinion, or knowledge of a product, service, or process. The responses are recorded and interpreted. The most common survey examples are written questionnaires, face-to-face or telephonic conversations, focus groups, and electronic (e-mail or...
9.5K
