用基于XLNet的方法进行软件漏洞检测的远程上下文建模,使用基于XLNet的方法
Yinhu Zhao1, Guanjun Lin2,3, Zhenxuan Liao4,5
1School of Electronic, Electrical Engineering and Physics, Fujian University of Technology, Fuzhou, Fujian, 350118, China.
Scientific reports
|January 16, 2026
概括
XLNetVD通过使用XLNet来捕获长代码依赖性来增强软件漏洞检测,以68%的F1得分优于现有模型. 这个框架提供了一个最先进的解决方案,用于识别代码中的漏洞.
科学领域:
- 网络安全 网络安全
- 软件工程 软件工程 软件工程
- 人工智能的人工智能
背景情况:
- 软件漏洞检测对于网络安全至关重要.
- 基于语言模型 (LM) 的方法显示出希望,但与远程代码依赖性作斗争.
- 变压器架构在捕获广泛的代码上下文方面存在局限性.
研究的目的:
- 介绍XLNetVD,这是一个基于XLNet的框架,用于功能级别的漏洞检测.
- 解决现有模型在捕获远程代码依赖性方面的局限性.
- 评估XLNet在漏洞检测方面的有效性.
主要方法:
- 利用双向变压器-XL模型进行扩展的上下文建模.
- 与六个上下文和三个非上下文嵌入模型对比的XLNet.
- 将XLNet集成到一个端到端的框架中,XLNetVD,并应用低级调整 (LoRA) 微调.
主要成果:
- XLNet获得了最高的F1得分68%,超过了CodeBERT和GPT-2.
- 在LoRA增强的LM中,XLNet-LoRA显示出最佳的性能效率权衡.
- 在不平衡的真实世界和平衡的SARD数据集上,XLNetVD表现出了竞争力.
结论:
- XLNetVD 确立了自己作为软件漏洞检测的最先进的解决方案.
- 该框架有效地捕捉了必要的代码依赖性,用于识别微妙的漏洞.
- 基于XLNet的方法比现有的基于LM的方法提供了显著的改进.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
