大型语言模型从整合跨位置束的指数窗口过渡到结构束的权力法窗口
David Skrill1, Sam V Norman-Haignere2,3
1Department of Biostatistics and Computational Biology, University of Rochester Medical Center, Rochester, NY 14642.
Advances in neural information processing systems
|March 4, 2024
概括
大型语言模型 (LLM) 开发特定的"集成窗口"来处理语言,反映大脑功能. 这些窗口从指数转移到跨网络层的结构约束的权力规律动态.
科学领域:
- 计算神经科学是一种计算神经科学.
- 人工智能的人工智能是人工智能.
- 自然语言处理自然语言处理.
背景情况:
- 现代语言模型 (LLM) 在处理语言方面与生物神经系统有相似之处.
- 人类大脑对语言的反应显示出分层的"整合窗口",限制了代币的影响.
- 有限的研究已经探索了LLMs内部的整合窗口.
研究的目的:
- 开发一种方法来估计黑盒LLMs的集成窗口.
- 在训练有素的法学士中描述时间整合模式.
- 调查整合窗口如何与结构性语言界限保持一致.
主要方法:
- 使用一种新的文字交换程序来估计没有梯度访问的LLM的集成窗口.
- 集成窗口的动态被模拟使用指数函数和权力定律函数的组合.
- 开发了一个指标来量化集成窗口和结构边界 (例如句子) 之间的关系.
主要成果:
- 训练有素的LLM展示了刻板印象的整合窗口,从指数转变为跨层的权力法动态.
- 集成窗口在后来的网络层中越来越多地与结构边界同步.
- 一个未经训练的模型显示了统一的整合,与训练的模型不同.
结论:
- 课程学习者学习自然语言的时间整合的刻板印象模式.
- 早期的层使用位置约束的指数窗口,而后来的层使用结构约束的权力规律窗口.
- 开发的方法提供了一个工具包,用于分析LLMs的时间整合,促进跨学科研究.
更多相关视频
相关概念视频
Improving Translational Accuracy
10.4K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.4K
Per-Unit Sequence Models
74
An ideal Y-Y transformer, grounded through neutral impedances, displays per-unit sequence networks akin to those of a single-phase ideal transformer when subjected to balanced positive- or negative-sequence currents. These currents do not produce neutral currents, and their associated voltage drops.
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...
74
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
498
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
498
Parametric Survival Analysis: Weibull and Exponential Methods
429
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
429
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
69
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
69
End Point Prediction: Gran Plot
324
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
324


