GATmath和GATLc:评估阿拉伯大语言模型的综合基准
Safa AlBallaa1, Nora AlTwairesh1, Abdulmalik AlSalman1
1Department of Computer Science, College of Computer and Information Sciences, King Saud University, Riyadh, Saudi Arabia.
开发阿拉伯大型语言模型 (LLM) 是一个挑战, 新的数据集GATmath和GATLc提供了大规模的推理和语言任务,以推动阿拉伯人工智能的发展.
科学领域:
- 人工智能
- 自然语言处理
- 计算语言学
背景情况:
- 大型语言模型 (LLM) 已经推进了人工智能,但它们的发展需要强有力的评估.
- 缺乏全面的基准和评估工具阻碍了对阿拉伯语LLM的评估.
- 这种稀缺性限制了阿拉伯语言模型的进步和实际应用.
研究的目的:
- 介绍GATmath (7千个问题) 和GATLc (9千个问题),这是阿拉伯语中多任务推理和语言理解的新标准.
- 提供首个专门为阿拉伯语言设计的大规模,全面的推理数据集.
- 促进严格的评估,并推动阿拉伯LLM的发展.
主要方法:
- 创建了两个大规模的阿拉伯语数据集,GATmath和GATLc,来自一般能力测试 (GAT).
- 数据集包括各种类别,需要推理,语义分析,语言理解和数学问题解决.
- 在这些新开发的基准上评估了七个著名的LLM.
主要成果:
- 最高效的LLM只达到66.9% (GATmath) 和64.3% (GATLc) 的准确性.
- 这些结果凸显了GATmath和GATLc数据集所带来的重大困难.
- 目前最先进的LLM在阿拉伯语推理和语言理解方面存在重大局限性.
结论:
- 对于现有的阿拉伯LLM来说,GATmath和GATLc数据集是一个相当大的挑战.
- 在开发更有能力的阿拉伯语言模型方面还有很大的改进空间.
- 这些基准对于推进阿拉伯人工智能研究和开发至关重要.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
相关概念视频
Improving Translational Accuracy
Multiple Comparison Tests
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
