轻量级语言模型对于复杂的计算表型化任务容易产生推理错误
Shashank Yadav1, David Maughan1, Vignesh Subbian1
1College of Engineering, The University of Arizona, Tucson, AZ.
ArXiv
|August 6, 2025
概括
大型语言模型 (LLM) 在复杂的计算表型化任务中显示推理错误. 加强像PHEONA这样的LLM评估框架对于识别和解决人工智能开发中的这些错误至关重要.
科学领域:
- 生物医学信息学 生物医学信息学
- 人工智能的人工智能
背景情况:
- 计算表型化对于队列识别至关重要,但由于手动数据审查,需要大量的时间.
- 之前的研究表明,LLM在复杂的表型化任务中存在局限性,特别是在多种疗法中.
研究的目的:
- 评估轻量级LLM在计算表型化中的推理能力.
- 加强PHEONA框架,用于评估LLMs中的错误推理.
主要方法:
- 评估了三种轻量级的LLM (DeepSeek-r1,Mistral Small,Phi-4) 进行表型准确性.
- 使用快速修改来识别解释正确性和不忠错误.
- 扩展了PHEONA框架,包括错误推理评估.
主要成果:
- 在所有测试的LLMs中,推理错误,包括解释的正确性和不忠诚性,普遍存在.
- 与Mistral和Phi相比,DeepSeek在快速修改后表现出最小的准确性影响.
- 增强的PHEONA框架成功地发现了普遍存在的推理错误.
结论:
- 推理错误在LLM对复杂任务的响应中无处不在,例如计算表型化.
- 增强的PHEONA框架对于LLM评估至关重要,强调需要改进可解释性方法.
相关概念视频
Types of Errors: Detection and Minimization
2.4K
Error is the deviation of the obtained result from the true, expected value or the estimated central value. Errors are expressed in absolute or relative terms.
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
Absolute error in a measurement is the numerical difference from the true or central value. Relative error is the ratio between absolute error and the true or central value, expressed as a percentage.
Errors can be classified by source, magnitude, and sign. There are three types of errors: systematic, random, and gross.
Systematic or...
2.4K
Mechanistic Models: Compartment Models in Individual and Population Analysis
86
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
86
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
101
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
101
Improving Translational Accuracy
11.9K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.9K
Errors In Hypothesis Tests
4.5K
When performing a hypothesis test, there are four possible outcomes depending on the actual truth (or falseness) of the null hypothesis and the decision to reject or not.
4.5K
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K


