Related Experiment Video
Updated: Jan 31, 2026

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
Filling the Gaps in Health Data: Using a Machine Learning Approach to Augment Partially Observed Variables Such as
Stefan Franzen1, Evangelos Chandakas2, Sam Hillman3
1BPM Evidence Statistics, AstraZeneca, Gothenburg, Sweden.
Purpose:
Missing information is common in real-world claims data, particularly on behavioral confounders, for example, smoking. Often one category of the variable, "yes" is partially observed while the other "no" remains completely missing-a pattern we call missing with truncation. A common way to handle these missing values is to naïvely treat missing values as absence of the risk factor, which may lead to substantial misclassification. Standard multiple imputation is impossible as only one level of the variable is observed.
Methods:
A case study was conducted using data from the NOVELTY study, including 12 224 people with physician diagnosed asthma and/or COPD (NCT02760329). From this cohort, 9733 patients with complete information were included. This dataset was split into two where the first part was used to train an imputation model and the second part was used to evaluate the imputations based on the model (1) when used to impute a truncated and amputated smoking variable against the naïvely classifying missing as "no" (2) when varying the percent smokers retained, q.
Results:
The accuracy of approaches (1) and (2) was 0.79 and 0.43, respectively; for q = 90%, the accuracy of approaches (1) and (2) was 0.89 and 0.94, respectively. Transfer learning showed better accuracy than the naïve approach when the percentage of true smokers being recorded as smokers was < 80%.
Conclusions:
The added value of transfer learning was greatest when low proportions of true ever-smokers were recorded, with its advantage depending on both the true prevalence of true smokers and the predictive model's performance.
More Related Videos
Related Concept Videos
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Data Reporting and Recording
Observational Learning
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...
Model Approaches for Pharmacokinetic Data: Compartment Models
Two primary types of compartment models are recognized: mammillary and catenary. The more...

