一个多功能数据集用于内在窃检测,文本重复使用分析和乌尔都语作者聚类
Muhammad Haseeb1, Muhammad Faraz Manzoor1, Muhammad Shoaib Farooq1
1Department of Computer Science, University of Management and Technology, Lahore, Pakistan.
Data in brief
|January 1, 2024
概括
这项研究引入了乌尔都语内在抄袭检测的新基准库,解决了对这种资源较少的语言缺乏资源的问题. 该数据集有助于开发精确的乌尔都文本抄袭检测模型.
科学领域:
- 自然语言处理 (NLP) 是一种自然语言处理.
- 计算语言学 计算语言学
- 信息检索 信息检索
背景情况:
- 盗版检测 (PD) 对学术诚信和知识产权至关重要.
- 现有的PD研究对高资源语言来说是广泛的,但对乌尔都语来说是有限的,特别是在内在抄袭检测方面.
- 乌尔都语是一个资源较低的语言,缺乏高质量的基准库来检测内在的抄袭行为.
研究的目的:
- 为解决乌尔都语内在抄袭检测资源的短缺问题.
- 为乌尔都语提供一种新的,高质量的基准语库.
- 促进乌尔都语专用PD模型的开发.
主要方法:
- 创建一个包含10872份乌尔都文档的基准库.
- 在句子和段落细节级别上构建语料库.
- 该集体旨在支持内在抄袭检测,字面文本重复使用识别和作者聚类.
主要成果:
- 已经建立了一个全面的,高质量的基准库,用于乌尔都语内在抄袭的检测.
- 该集体支持多个NLP任务,包括内在PD,文本重用和作者聚类.
- 该数据集填补了乌尔都语NLP研究资源的关键缺口.
结论:
- 开发的乌尔都语语库对NLP研究做出了重大贡献,特别是在低资源语言方面.
- 这个资源将使得乌尔都语更准确,更有效的抄袭检测系统的创建成为可能.
- 该集体将提高乌尔都语内容的教育和出版部门的抄袭检测能力.
相关概念视频
Proteomics
7.3K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.3K
DNA Isolation
193.0K
DNA from cells is required for many biotechnology and research applications, such as molecular cloning. To remove and purify DNA from cells, researchers use various methods of DNA extraction. While the specifics of different protocols may vary, some general concepts underlie the process of DNA extraction.
193.0K
Data Collection I
6.2K
Data collection gathers information needed to make accurate judgments about a patient's present condition. During a health history interview, subjective data is collected from the patient, their caregivers, or family members, and objective data is collected through observations and physical assessment. Patients are the primary source of subjective data. Thus information gathered from patients through interviews, observations, and physical examination is primary data. Secondary sources of...
6.2K
Proofreading
6.3K
Synthesis of new DNA molecules is carried out by the enzyme DNA polymerase, which adds nucleotides on the daughter strand complementary to the template DNA strand. DNA polymerase has a higher affinity to add the correct base and ensures fidelity during DNA replication. Furthermore, it exhibits proofreading activity during replication, using an exonuclease domain that cuts off incorrect nucleotides from the nascent DNA strand.
Errors During Replication are Corrected by the DNA Polymerase...
Errors During Replication are Corrected by the DNA Polymerase...
6.3K
Data Collection II
8.2K
The nursing history captures and records the patient's health status, so that a care plan evolves to meet the patient's individual needs. The nursing health history is a part of the initial assessment. A comprehensive history covers all health dimensions and plays a significant role in the assessment process. A comprehensive history includes the patient's biographical information, reasons for seeking health care, expectations, present and past health history, medications, and...
8.2K
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K


