通过扩散模型可靠地生成保护隐私的合成电子健康记录时间序列
Muhang Tian1, Bernie Chen2, Allan Guo1
1Department of Computer Science, Duke University, Durham, NC 27708, United States.
Journal of the American Medical Informatics Association : JAMIA
|September 2, 2024
概括
这项研究引入了一种新的扩散模型,用于生成现实的,保护隐私的合成电子健康记录 (EHR). 该方法提高了医疗研究的数据实用性,同时减轻了与传统身份删除技术相关的隐私风险.
科学领域:
- 医疗信息学 医疗信息学
- 机器学习 机器学习
- 数据 隐私 数据 隐私 数据
背景情况:
- 电子健康记录 (EHR) 对医疗数据分析有价值,但面临隐私限制.
- 现有的去识别方法有缺陷,导致隐私泄露和有限的数据可用性.
- 推进医学研究需要克服EHR访问障碍.
研究的目的:
- 开发一种方法来生成现实的和保护隐私的合成EHR时间序列数据.
- 为了解决当前EHR去识别和研究数据可用性的局限性.
- 通过改善数据访问,促进医疗保健中的机器学习应用.
主要方法:
- 利用无声扩散概率模型生成合成EHR时间序列.
- 在多个医疗保健数据库 (MIMIC-III,MIMIC-IV,eICU) 和非EHR数据集上进行了实验.
- 将拟议的方法与八种现有的非识别和数据生成技术进行了比较.
主要成果:
- 拟议的扩散模型在数据保真方面明显优于现有方法.
- 生成的合成数据的分辨准确度较低,表明隐私风险降低.
- 与基线方法相比,该方法需要较少的培训工作.
结论:
- 扩散模型产生现实的合成电子健康记录,保护患者的隐私.
- 这种方法可以缓解医疗保健中的数据可用性问题,减少EHR访问的障碍.
- 该方法支持用于健康研究的机器学习的进步.
更多相关视频
11:21Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
8.2K
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
10.7K
相关概念视频
Introduction To Survival Analysis
197
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
197
Mechanistic Models: Compartment Models in Individual and Population Analysis
33
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
33
