在共享临床数据集中重新识别参与者:实验研究实验研究
Daniela Wiepert1, Bradley A Malin2,3,4, Joseph R Duffy1
1Department of Neurology, Mayo Clinic, Rochester, MN, United States.
JMIR AI
|June 14, 2024
概括
语音录音的重新识别风险较低,特别是在更大的数据集和非连接的语音任务中. 这项研究为共享临床语音数据提供了隐私措施的信息.
科学领域:
- 生物识别信息 生物识别信息
- 语音处理 语音处理
- 医疗信息学 医疗信息学
背景情况:
- 大量的语音数据集对于医疗保健工具至关重要.
- 分享语音记录引发了隐私问题,特别是在受保护的健康信息方面.
- 现有的扬声器识别方法存在潜在的重新识别风险.
研究的目的:
- 评估从语音录音中重新识别个人的风险.
- 评估搜索空间大小和语音任务类型对重新识别风险的影响.
- 为临床语音数据共享提供隐私保护策略的信息.
主要方法:
- 使用最先进的扬声器识别模型建模了一个对抗性攻击.
- 使用VoxCeleb.测试了各种数据集大小 (已知和未知数据集) 的重新识别风险.
- 研究了不同语音录音类型 (连接与非连接) 对使用临床数据集重新识别的影响.
主要成果:
- 随着搜索空间 (比较数量) 的增加,重新识别风险会降低.
- 在错误的接受和比较数量之间观察到正线性相关性.
- 非连接的语音任务 (例如,母音延长) 与交叉任务条件下的连接语音任务相比,具有较低的重新识别风险.
结论:
- 扬声器识别模型在某些条件下可以重新识别参与者,但实际风险似乎很小.
- 搜索空间大小和语音任务类型显著影响重新识别风险.
- 调查结果提供了可行的建议,以加强参与者在语音数据共享和政策制定中的隐私.
相关概念视频
Crossover Experiments
2.8K
Crossover experiments, also called the repeated-measurements design, is a study design in which all experimental units are exposed to all treatments in different periods. Crossover experiments are generally used in psychology, the pharmaceutical industry, agriculture, and medicine.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
2.8K
Blinding
2.4K
Blinding is a commonly used method of not telling participants which treatment a subject is receiving. Blinding is a critical part of a randomized control trial or RCT. It reduces the bias that affects the results. In an RCT, blinding is used in the form of a placebo. A placebo effect occurs when untreated subjects falsely believe they have received the treatment and report improved symptoms. A placebo or a dummy treatment is administered to subjects to negate the bias caused by such an effect.
2.4K
Data Collection by Experiments
24.1K
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public...
An example of the experimental method is a public...
24.1K
Randomized Experiments
6.9K
The randomization process involves assigning study participants randomly to experimental or control groups based on their probability of being equally assigned. Randomization is meant to eliminate selection bias and balance known and unknown confounding factors so that the control group is similar to the treatment group as much as possible. A computer program and a random number generator can be used to assign participants to groups in a way that minimizes bias.
Simple randomization
Simple...
Simple randomization
Simple...
6.9K
Blind Procedures
10.6K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
10.6K
Longitudinal Research
11.9K
Sometimes we want to see how people change over time, as in studies of human development and lifespan. When we test the same group of individuals repeatedly over an extended period of time, we are conducting longitudinal research. Longitudinal research is a research design in which data-gathering is administered repeatedly over an extended period of time. For example, we may survey a group of individuals about their dietary habits at age 20, retest them a decade later at age 30, and then again...
11.9K


