增强基于瘤学试验匹配的生物标志物,使用大型语言模型.
Nour Alkhoury1, Maqsood Shaik1, Ricardo Wurmus1
1Berlin Institute for Medical Systems Biology (BIMSB), Max Delbrück Center for Molecular Medicine, Berlin, Germany.
NPJ digital medicine
|May 5, 2025
概括
开源语言模型有效地从临床试验数据中提取基因组生物标志物,优于闭源替代品. 微调进一步提高了他们对癌症药物开发的关键信息进行结构化的能力.
科学领域:
- 在瘤学瘤学.
- 生物信息学是一种生物信息学.
- 自然语言处理自然语言处理.
背景情况:
- 临床试验的招生依赖于患者的资格标准,通常在非结构化文本中发现.
- 基因组生物标志物对于精准医学和向癌症治疗至关重要.
- 将患者与临床试验相匹配需要有效地提取资格信息.
研究的目的:
- 探索从瘤学临床试验描述中提取遗传生物标志物的策略.
- 结构化非结构化临床试验数据,以改善患者匹配.
- 评估大型语言模型 (LLM) 在这个任务中的表现.
主要方法:
- 利用开源和闭源大语言模型 (LLM) 来处理临床试验研究描述.
- 专注于从资格标准中提取和结构化基因组生物标记信息.
- 将准备好的LLM性能与精心调整的模型进行比较.
主要成果:
- 开源的LLM在捕获复杂的逻辑表达式和结构化基因组生物标志物方面表现出有效性.
- 开放源代码模型的性能优于像GPT-4这样的封闭源代码模型.
- 用额外数据微调开源模型导致了显著的性能提升.
结论:
- 开源的LLM是一种可行的解决方案,用于结构化非结构化临床试验数据,特别是基因组生物标志物.
- 这些模型有助于识别适合精确瘤学试验的患者.
- 进一步开发和微调可以优化临床试验数据提取的LLM性能.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


