Related Experiment Video
Updated: Aug 16, 2025

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Large-scale Vietnamese point-of-interest classification using weak labeling.
Van Trung Tran1,2, Quang Dao Le1,2, Bao Son Pham3
1Center of Multidisciplinary Integrated Technologies for Field Monitoring, Vietnam National University of Engineering and Technology, Hanoi, Vietnam.
This study introduces a large Vietnamese Point-of-Interests (POI) dataset for classification. Weak labeling significantly improved POI data accuracy, achieving a 90% F1 score.
Area of Science:
- Computer Science
- Geographic Information Systems
- Natural Language Processing
Background:
- Crowd-sourced Point-of-Interests (POI) data often suffers from low-quality category labels, impacting location-based applications.
- Existing Vietnamese POI datasets are limited in size and scope, hindering advancements in localized digital mapping.
- Accurate POI categorization is crucial for services ranging from navigation to local business discovery.
Purpose of the Study:
- To create the first large-scale, annotated dataset for categorical classification of Vietnamese Points of Interest (POIs).
- To develop and validate a weak labeling approach for efficient POI dataset creation.
- To establish a benchmark for POI categorical classification in the Vietnamese language domain.
Main Methods:
- Collected 750,000 POIs from a Vietnamese digital map (WeMap).
- Implemented a novel weak labeling strategy to overcome the limitations of manual annotation.
- Constructed a dataset comprising 275,000 weak-labeled POIs for training and 30,000 gold-standard POIs for testing.
- Utilized BERT-based fine-tuning as a strong baseline for empirical evaluation.
Main Results:
- The developed dataset is the largest of its kind for Vietnamese POIs, covering 15 categories.
- The weak labeling approach proved highly efficient for large-scale data annotation.
- The BERT-based baseline achieved a 90% F1 score on the gold-standard test set.
- Significant improvement in WeMap's POI data accuracy, increasing from 56% to 93%.
Conclusions:
- The proposed weak labeling method is effective for creating large-scale, high-quality POI datasets.
- The new dataset provides a valuable resource for advancing POI classification research in Vietnamese.
- The findings demonstrate the feasibility and scalability of automated POI categorization using deep learning models.
Related Concept Videos
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Classification of Systems-II
Methods of Classification and Identification
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...

