Related Experiment Video
Updated: Jul 2, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.5K
NERSkill.Id: Annotated dataset of Indonesian's skill entity recognition
Meilany Nonsi Tentua1, Suprapto2, Afiahayati2
1Informatic, Sains and Technology Faculty, Universitas PGRI Yogyakarta, Indonesia.
Data in Brief
|February 26, 2024
Summary
NERSkill.Id is a new dataset for Indonesian skill recognition, identifying hard skills, soft skills, and technology. This resource aids NLP research and improves talent acquisition by enabling accurate skill identification.
Area of Science:
- Natural Language Processing
- Computational Linguistics
- Data Science
Background:
- Annotated datasets are crucial for training Natural Language Processing (NLP) models.
- Existing annotated corpora for the Indonesian language, particularly for skill entities, are scarce.
- Accurate skill identification is vital for talent acquisition, job matching, and educational program development.
Purpose of the Study:
- To introduce NERSkill.Id, a manually annotated Named Entity Recognition (NER) dataset for skill entities in Indonesian.
- To provide a valuable resource for advancing NLP research in Indonesian.
- To support stakeholders in improving talent acquisition and educational alignment with market needs.
Main Methods:
- Data collection from a job portal.
- Manual annotation of tokens following the BIO (Beginning, Inside, Outside) scheme.
- Processing using open-source libraries for dataset construction.
Main Results:
- The NERSkill.Id dataset contains 418,868 tokens.
- 15.51% of tokens are identified as named entities, categorized into hard skill, soft skill, and technology.
- The dataset is suitable for training and evaluating NER systems for skill extraction.
Conclusions:
- NERSkill.Id addresses the scarcity of annotated Indonesian corpora for skill entities.
- The dataset facilitates advancements in skill entity recognition for the Indonesian language.
- It offers significant benefits for NLP researchers, companies, recruiters, and educational institutions.
Keywords:
Indonesian skill entityNamed entity recognitionNatural language processingSkill entity recognitionText miningMore Related Videos
Related Concept Videos
Classification of Signals
461
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
461
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Data Collection by Observations
12.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
12.0K
Stratified Sampling Method
12.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.0K
RACE - Rapid Amplification of cDNA Ends
6.3K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.3K
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K

