相关实验视频
零射击 低级专家的稀少混合物 从预训练的基础模型构建预训练的基础模型
IEEE transactions on pattern analysis and machine intelligence
|September 22, 2025
概括
深度模型融合利用预训练模型加速开发. 低级专家稀少混合 (SMILE) 结构可以实现高效的升级,显著减少参数干扰,并通过最小的额外参数提高性能.
科学领域:
- 人工智能的人工智能
- 机器学习 机器学习
- 深度学习 (Deep Learning) 是一种深度学习.
背景情况:
- 在广泛的数据集上进行深度模型培训是昂贵的.
- 深度模型融合通过利用先前存在的模型提供了一个解决方案.
- 参数干扰和缺乏可解释性是模型融合中的关键挑战.
研究的目的:
- 为了解决深度模型融合中的参数干扰.
- 引入一种高效的方法,将源模型扩展到专家混合 (MoE) 模型中.
- 为了提高模型性能,加快新的模型开发,而无需额外的培训数据.
主要方法:
- 使用子空间分析检查了线性层的微调.
- 定义参数干扰作为一个优化问题.
- 引入了低级专家零射击SparseMIxture (SMILE) 结构,用于模型升级.
主要成果:
- 在没有额外的数据或培训的情况下,SMILE将升级级型号纳入MoE架构.
- 维度扩展有效地管理参数干扰.
- 在8个单独的ViT车型中实现了98%-99%的性能,并增加了50%的额外参数以实现完整的微调.
- 对于LoRA微调的Flan-T5型号,性能保持在99%,仅有2%的额外参数.
结论:
- 微笑在图像分类,文本生成和大型语言模型 (LLM) 中展示了适应性和可扩展性.
- 该方法有效地利用预训练模型的知识,克服参数干扰的挑战.
- 微笑为深度模型开发和融合提供了一种具有成本效益的方法.
相关概念视频
Residuals and Least-Squares Property
9.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
9.1K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K
Improving Translational Accuracy
3.6K
3.6K
Upsampling
588
Managing signal sampling rates is essential in digital signal processing to maintain signal integrity. A decimated signal, characterized by a reduced frequency range due to its lower sampling rate, can be upsampled by inserting zeros between each sample. This upsampling process expands the original spectrum and introduces repeated spectral replicas at intervals dictated by the new Nyquist frequency. To refine this zero-inserted sequence, it is passed through a lowpass filter with a cutoff...
588
Extraction: Advanced Methods
1.1K
Metal ions can be separated from one another by complexation with organic ligands–the chelating agent– to form uncharged chelates. Here, the chelating agent must contain hydrophobic groups and behave as a weak acid, losing a proton to bind with the metal. Since most organic ligands used in this process are insoluble or undergo oxidation in the aqueous phase, the chelating agent is initially added to the organic phase and extracted into the aqueous phase. The metal-ligand complex is...
1.1K
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
1.1K
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
1.1K