临床前基因组基准:用于评估临床前基因病理分类的大型语言模型的试点基准数据集
Avan Kader1, Marie-Luise H H Ranner-Hafferl2, Felix Reuter1
1Department of Diagnostic and Interventional Radiology, Technical University of Munich, Ismaninger Str. 22, 81675 Munich, Germany.
Biology
|March 13, 2026
概括
本研究引入了一个基准数据集,用于评估临床前病理学中的大语言模型 (LLM). 法学士的成绩差异很大,显示出对阶级不平衡的敏感性和作为研究选工具的潜力.
科学领域:
- * 临床前组织病理学
- * 在病理学中的人工智能
- * 大型语言模型 (LLM) 评估.
背景情况:
- *缺乏标准化的基准来评估临床前病理学LLM.
- *需要在病理学AI模型中使用多维分类能力.
- * 开发一个试点数据集,用于LLM评估在这个领域.
研究的目的:
- * 创建和呈现一个基准数据集,用于评估LLM在组织样本上的表现.
- *对包括物种,器官,染色和制剂类型在内的多维分类任务进行LLM的评估.
- * 解决在人工智能驱动的基因病理学中对标准化评估指标的需求.
主要方法:
- * 在378个临床前组织学样本上对三种LLM (GPT-4.1,GPT-4o-mini,Llama 3.2) 的评估.
- *四个分类维度:物种 (老鼠,子,老鼠),器官,染色方法和制剂类型 (冷或嵌入).
- *使用灵敏度,特异性和混矩阵分析进行绩效评估.
主要成果:
- *跨任务的LLM绩效存在实质性的差异,对阶级不平衡的高度敏感.
- * GPT-4.1显示了平衡的制剂类型分类;Llama 3.2在抛样本方面遇到了困难.
- *拉玛3.2识别了所有物种,但对老鼠的识别能力很差;GPT-4.1在识别老鼠方面表现出色.
- * Llama 3.2 显示了高染色分类性能; GPT-4o-mini 实现了完美的 H&E 识别.
结论:
- *目前的LLM在组织学分类中表现出可变的性能,对类不平衡非常敏感.
- *LLM不适合在组织病理学中独立诊断使用.
- * 在人类监督下,LLM在研究环境中显示出选工具的潜力.
更多相关视频
09:06Whole-brain Segmentation and Change-point Analysis of Anatomical Brain MRI—Application in Premanifest Huntington's Disease
Published on: June 9, 2018
12.7K
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
8.9K
相关概念视频
Genetic Lingo
Overview
Genome-wide Association Studies-GWAS
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
