Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Quantifying and Rejecting Outliers: The Grubbs Test01:02

Quantifying and Rejecting Outliers: The Grubbs Test

1.6K
Sometimes, a data set can have a recorded numerical observation that greatly  deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier.  To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Econometric Views (EViews)01:29

Econometric Views (EViews)

145
Econometric Views, often stylized as EViews, is a package that merges statistical analysis with econometric studies. It is designed to provide tools for time series analysis, forecasting, and econometric model simulation. The software originated from MicroTSP software and has evolved significantly since its inception in 1981. The history of EViews is marked by a continuous effort to enhance its computational speed and user interface. It was initially developed for large computing systems but...
145
Outliers and Influential Points01:08

Outliers and Influential Points

4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Regression Analysis01:11

Regression Analysis

5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Aggregates Classification01:29

Aggregates Classification

326
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
326
Prediction Intervals01:03

Prediction Intervals

2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.3K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Data of dwelling time process at container terminal: Multi-perspective based dataset.

Data in brief·2024
Same author

Data set for Gambung green tea aroma using on electronic nose.

BMC research notes·2024
Same author

Machine learning approach for predicting production delays: a quarry company case study.

Journal of big data·2022
Same author

Electronic nose homogeneous data sets for beef quality classification and microbial population prediction.

BMC research notes·2022
Same author

Electronic nose dataset for pork adulteration in beef.

Data in brief·2020
Same author

Electronic nose dataset for beef quality monitoring in uncontrolled ambient conditions.

Data in brief·2018

Related Experiment Video

Updated: Jul 4, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.5K

Poverty prediction using E-commerce dataset and filter-based feature selection approach.

Dedy Rahman Wijaya1, Raden Ilham Fadhilah Ibadurrohman2, Elis Hernawati2

  • 1School of Applied Science, Telkom University, Bandung, Indonesia. dedyrw@telkomuniversity.ac.id.

Scientific Reports
|February 6, 2024
PubMed
Summary

This study proposes using e-commerce data and machine learning to predict poverty levels faster than traditional surveys. The best results combined f-score feature selection with support vector regression for accurate poverty rate prediction.

More Related Videos

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

727
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K

Related Experiment Videos

Last Updated: Jul 4, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
07:35

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances

Published on: October 11, 2018

7.5K
Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
03:37

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers

Published on: March 1, 2024

727
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
07:15

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model

Published on: August 16, 2020

6.8K

Area of Science:

  • Socioeconomics
  • Data Science
  • Computational Social Science

Background:

  • Poverty assessment traditionally relies on time-consuming surveys and censuses.
  • Governments require rapid insights into socio-economic conditions for effective development planning.
  • Existing methods for poverty data collection are resource-intensive and slow.

Purpose of the Study:

  • To develop a faster proxy for poverty level assessment using e-commerce data and machine learning.
  • To investigate the efficacy of various machine learning algorithms in predicting poverty rates.
  • To propose a novel approach combining feature selection with predictive modeling for poverty mapping.

Main Methods:

  • Utilized a high-dimensional e-commerce dataset for poverty prediction.
  • Employed statistical-based feature selection algorithms (e.g., f-score) to identify relevant predictors.
  • Compared three machine learning models: support vector regression, linear regression, and k-nearest neighbor.

Main Results:

  • The combination of f-score feature selection and support vector regression demonstrated superior performance in predicting poverty rates.
  • E-commerce data, when processed with appropriate feature selection, proved effective for poverty level estimation.
  • The proposed methodology offers a significant speed advantage over conventional survey-based approaches.

Conclusions:

  • E-commerce data and machine learning algorithms present a viable and efficient proxy for poverty level prediction.
  • The integration of feature selection enhances the accuracy and efficiency of machine learning models for socio-economic analysis.
  • This approach can provide timely information for policymakers to inform area development strategies.