填补健康数据的空白:使用机器学习方法来增加部分观察到的变量,例如在索赔数据中吸烟
Stefan Franzen1, Evangelos Chandakas2, Sam Hillman3
1BPM Evidence Statistics, AstraZeneca, Gothenburg, Sweden.
Pharmacoepidemiology and drug safety
|January 30, 2026
概括
转移学习有效地将声称中缺少的吸烟数据归因于声称,当记录的吸烟者很少时,优于天真的方法. 这提高了健康研究中行为混分析的准确性.
科学领域:
- 医疗信息学 医疗信息学
- 生物统计学 生物统计学
- 流行病学 流行病学
背景情况:
- 现实世界索赔数据往往缺乏行为混因素,如吸烟状况.
- 一个特定的模式",缺失与截断",发生在"是"部分观察,但"不"完全缺失.
- 轻率地将缺少的吸烟数据视为"不"可以导致严重的错误分类.
研究的目的:
- 评估转移学习,以在健康声明中赋予截断的吸烟数据.
- 将转移学习与对待缺少数据作为没有风险的天真方法进行比较.
- 在观察到的吸烟者的不同比例下评估归算准确性.
主要方法:
- 一个案例研究使用了来自NOVELTY研究 (NCT02760329) 的数据,其中包括9733名患者.
- 在一个数据子集上训练了一个归算模型,并在另一个数据子集上进行评估.
- 使用转移学习与天真的"错过等于没有"方法比较模型的性能,并改变保留的吸烟者的百分比 (q).
主要成果:
- 与天真方法 (0.79) 相比,转移学习实现了更高的准确性 (0.89),用于归因吸烟状态.
- 当90%的吸烟者被保留时 (q=90%),转移学习的准确性达到0.94,而天真方法的准确性为0.89.
- 转移学习在记录不到80%的真正吸烟者时表现出更高的准确性.
结论:
- 转移学习为赋予吸烟数据提供了显著的附加价值,特别是当记录的真正常吸烟者很少时.
- 转移学习的好处取决于吸烟的真实流行率和预测模型的准确性.
- 这种方法增强了在大型健康数据集中处理缺失的行为混因素的能力.
更多相关视频
相关概念视频
Data Collection by Observations
15.0K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
15.0K
Data Reporting and Recording
5.4K
Reporting and recording are crucial in data documentation. The timely, thorough, and accurate documentation of facts is essential when recording patient data. Failure to record findings during an assessment or interpretation of a problem will result in loss of information and make the patient document unreliable. The reader is left with general impressions if the information is not specific. A recording is documenting data of the individual's health information in a traceable, secure, and...
5.4K
Observational Learning
969
Albert Bandura's observational learning, also known as imitation or modeling, occurs when a person observes and imitates another's behavior. It is a quicker process than operant conditioning. A well-known example is the Bobo doll study, where children who saw an adult acting aggressively towards the doll were more likely to act aggressively when left alone, compared to those who observed a nonaggressive adult. Many psychologists view observational learning as a form of latent learning...
969
How Data are Classified: Categorical Data
44.5K
A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
44.5K
How Data are Classified: Numerical Data
38.0K
Data that are countable or measurable in specific units are called numerical or quantitative data. Quantitative data are always numbers. Quantitative data are the result of counting or measuring the attributes of a population. Amount of money, pulse rate, weight, number of people living in a town, and number of students who opt for statistics are examples of quantitative data.
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
38.0K
Model Approaches for Pharmacokinetic Data: Compartment Models
554
Compartmental analysis is a widely adopted approach to characterizing drug pharmacokinetics. It uses compartment models that conceptualize the body as a collection of reversibly communicating compartments, each representing a group of tissues exhibiting similar drug distribution characteristics. The movement rate of the drug between these compartments is typically described by first-order kinetics.
Two primary types of compartment models are recognized: mammillary and catenary. The more...
Two primary types of compartment models are recognized: mammillary and catenary. The more...
554


