一个全器官视觉语言模型,用于可概括的3DCT表示
Cameron Beeche1,2, Joonghyun Kim1, Hamed Tavolinejad2,3
1Department of Bioengineering, University of Pennsylvania, Philadelphia, PA, 19104, USA.
medRxiv : the preprint server for health sciences
|July 9, 2025
概括
珀西瓦尔是一种新的AI基础模型,通过对各种数据进行训练,提高了计算机断层扫描 (CT) 医学成像的概括性. 这种人工智能工具提高了临床工作流程的效率,并揭示了疾病表型.
科学领域:
- 医疗成像中的人工智能
- 医疗保健基金会模型的基础
- 放射学人工智能的人工智能
背景情况:
- 现有的计算机断层扫描 (CT) 成像AI模型由于狭窄的培训环境而缺乏通用性.
- 有限的解剖覆盖,对比设置和临床指示限制了当前人工智能工具的现实应用.
- 存在对人工智能模型的需求,这些模型可以有效地处理广泛的体积CT成像数据.
研究的目的:
- 介绍Percival,一个新的视觉语言基础模型用于CT医学成像.
- 提高AI模型在临床放射学工作流程中的通用性.
- 评估帕西瓦尔发现的临床相关性和潜在结构.
主要方法:
- 开发了Percival,一个双编码器视觉语言基础模型.
- 培训了珀西瓦尔在超过40万张CT卷和来自宾夕法尼亚大学医学生物银行的配对放射学报告上.
- 使用基于变压器的图像编码器和BERT风格的语言编码器,通过对称对比学习对齐.
- 在超过20,000名参与者的成像数据 (100,000+ CT 卷) 上验证了 Percival.
主要成果:
- 与在有限数据上训练的模型相比,珀西瓦尔在图像-文本回忆任务中表现出卓越的表现.
- 通过关联研究和生存分析评估了珀西瓦尔的临床知识.
- 在数据中发现了丰富的潜在结构,与生理测量和疾病表型保持一致.
结论:
- 珀西瓦尔代表了计算机断层成像可通用的AI的重大进步.
- 该模型在多种CT数据中进行概括的能力提高了其提高临床工作流效率的潜力.
- 珀西瓦尔的分析揭示了对生物,表型和预后因素的临床相关见解.
相关概念视频
Vision
55.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
55.4K
Computed Tomography
6.3K
Tomography refers to imaging by sections. Computed tomography (CT) is a non-invasive imaging technique that uses computers to analyze several cross-sectional X-rays to reveal minute details about structures in the body.
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
6.3K
Depth Perception and Spatial Vision
952
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
952


