相关实验视频
用微调变压器和LLM模型在低资源环境中创建数据集并对克什米尔新闻片段分类进行基准测试
Deheem U Deyar1, Anirud Ramani1, Deepa Gupta2
1Department of Computer Science and EngineeringAmrita School of Computing, Amrita Vishwa Vidyapeetham, Bengaluru, India.
Scientific reports
|November 19, 2025
概括
本研究引入了用于自然语言处理 (NLP) 任务的新克什米尔新闻数据集. 微调的ParsBERT-Uncased在分类克什米尔新闻片段方面取得了很高的准确性.
科学领域:
- 自然语言处理 (NLP) 是一种自然语言处理.
- 计算语言学 计算语言学
- 低资源的语言技术
背景情况:
- 克什米里语是一个资源不足的语言,NLP数据集有限.
- 缺乏注释数据阻碍了卡什米尔语NLP应用程序的开发.
研究的目的:
- 通过创建一个标记的数据集来解决克什米尔NLP资源的稀缺问题.
- 为克什米尔新闻片段开发和评估有效的文本分类模型.
主要方法:
- 使用微软Bing将英语新闻片段翻译成克什米尔语.
- 手动细化和分类15,036个克什米尔新闻片段到十个领域.
- 机器学习,深度学习,变压器模型和大型语言模型 (LLM) 的实验.
主要成果:
- 创建了一个手动标记的数据集,包含10个类别的15,036个克什米尔新闻片段.
- 微调的ParsBERT-Uncased实现了最高的性能,F1得分为0.98.
- 确定了克什米尔语文本分类的有效方法.
结论:
- 创建的数据集是对克什米尔NLP的重要贡献.
- 这项研究为准确的克什米尔文本分类提供了一个强大的框架.
- 这项研究推进了卡什米里语等代表性不足的语言的NLP.
相关概念视频
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Aggregates Classification
956
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
956