NERSkill.Id:印尼技能实体认可的注释数据集
Meilany Nonsi Tentua1, Suprapto2, Afiahayati2
1Informatic, Sains and Technology Faculty, Universitas PGRI Yogyakarta, Indonesia.
Data in brief
|February 26, 2024
概括
NERSkill.Id是一个新的数据集,用于印尼技能识别,识别硬技能,软技能和技术. 该资源有助于NLP研究,并通过准确识别技能来改善人才获取.
科学领域:
- 自然语言处理自然语言处理.
- 计算语言学 计算语言学
- 数据科学数据科学数据科学
背景情况:
- 有注释的数据集对于训练自然语言处理 (NLP) 模型至关重要.
- 现有的印尼语注释体,特别是技能实体的注释体,很少.
- 准确的技能识别对于人才获取,工作匹配和教育计划开发至关重要.
研究的目的:
- 介绍NERSkill.Id,一个手动注释的命名实体识别 (NER) 数据集,以印尼语为技能实体.
- 为推进印尼语NLP研究提供有价值的资源.
- 支持利益相关者改善人才获取和教育与市场需求的调整.
主要方法:
- 从就业门户网站收集数据.
- 按照BIO (开始,内部,外部) 方案对代币进行手动注释.
- 使用开源库进行数据集构建的处理.
主要成果:
- 在NERSkill.Id数据集中包含418,868个令牌.
- 15.51%的代币被确定为有名称的实体,分为硬技能,软技能和技术.
- 该数据集适用于培训和评估NER系统以提取技能.
结论:
- NERSkill.Id解决了对技能实体的印尼注释体的稀缺问题.
- 该数据集促进了印尼语言技能实体认可的进步.
- 它为NLP研究人员,公司,招聘人员和教育机构提供了显著的好处.
更多相关视频
相关概念视频
Classification of Signals
461
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
461
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Data Collection by Observations
12.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
12.0K
Stratified Sampling Method
12.0K
Sampling is a technique to select a portion (or subset) of the larger population and study that portion (the sample) to gain information about the population. The sampling method ensures that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a stratified sample, divide the population into groups called strata and then take a...
To choose a stratified sample, divide the population into groups called strata and then take a...
12.0K
RACE - Rapid Amplification of cDNA Ends
6.3K
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific...
6.3K
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K


