在LLMs中对判决的模拟
Edoardo Loru1, Jacopo Nudo2, Niccolò Di Marco3
1Department of Computer, Control and Management Engineering, Sapienza University of Rome, Rome 00185, Italy.
概括
大型语言模型 (LLM) 与人类相比显示出不同的评估模式,依赖于词汇关联和统计先验. 这可能会导致"流行病症",一种知识的错觉,其中表面可信性取代了验证.
科学领域:
- 人工智能的人工智能
- 认知科学 认知科学
- 信息科学 信息科学 信息科学
背景情况:
- 大型语言模型 (LLM) 越来越多地用于评估任务,如可信度评估.
- 了解LLM评估策略至关重要,因为它们日益融入信息系统.
- 在LLM和人类评估机制之间的差异需要调查.
研究的目的:
- 在评估任务中将大型语言模型 (LLM) 与专家评级和人类判断进行基准测试.
- 分析指导LLM评估的基本机制和假设.
- 将LLM评估策略与人类推理过程进行比较.
主要方法:
- 为了直接比较,实施了一个结构化的代理框架.
- 六个LLM与NewsGuard和媒体偏见/事实检查专家评级进行了对比.
- 非专家的人类参与者遵循与LLM相同的评估程序 (标准选择,内容检索,理由).
主要成果:
- 尽管输出对齐,但LLM在可观测的评估标准中表现出一致的差异.
- 法学士评估似乎受到词汇关联和统计先验的影响,与人类的上下文推理有所不同.
- 观察到一种将语言形式与认识论可靠性混的倾向,称为"流行病学".
结论:
- 法学士评估策略可能与人类规范推理有很大差异,倾向于基于模式的近似推理.
- 将判断权委托给LLM可能会改变评估过程中的基本启发式.
- 该研究提出了关于LLMs在复杂的评估功能中的作用和影响的关键问题.
相关概念视频
Language and Cognition
711
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
711
Language Development
848
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
848
Modeling and Similitude
613
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
613
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.5K
3.5K
Prediction Intervals
3.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.3K

