Related Experiment Video
Updated: Aug 19, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Representational Veracity in Data Science Health Research: Targets, Proxies, Labels, and Descriptors
Clement Adebamowo1,2, Sally N Adebamowo1,2, Adeola Akintola3
1University of Maryland Marlene and Stewart Greenebaum Comprehensive Cancer Center, 660 W Redwood St., Baltimore, US.
Unstructured:
Prevailing ethical oversight of data science health research concentrates on privacy, consent, bias, and fairness. These concerns are necessary but insufficient, because each presupposes an answer to a prior question that is seldom asked directly. That is "do the targets, proxies, labels, classifications, ontologies, and population descriptors on which a current study rests still truthfully represent the persons, populations, and phenomena they are taken to describe, at the point of use rather than the point of collection?" In this viewpoint we name that question representational veracity (RV) and develop it as a construct for upstream ethical review. Our aims are to define RV and derive the domains along which it can be assessed; to demonstrate that it asks something that measurement validity, critical data studies, and algorithmic fairness do not; and to translate it into instruments that review bodies can use. We derive four assessment domains of RV analytically, asking for each transition in the data journey what must remain stable for a stored artifact still to stand for what it originally stood for. The resulting domains are material provenance, informational descriptors, normative authorization, and relational community. These domains interact but do not substitute for one another. Intact provenance cannot repair a poorly chosen target, and a transparent labeling process cannot confer authorization it never had. Drawing on scholarship in quantification, classification, measurement, critical data studies, algorithmic fairness, and health artificial intelligence governance, we show that a model may be accurate, reproducible, and formally fair while resting on a representation that is too thin, too unstable, or too normatively misdirected for the proposed use. We examine four recurrent failure modes, proxy substitution, category misassignment, label generation error, and descriptor sedimentation, anchoring each in a published case, and we present a counterpoint in which better representation reveals rather than conceals inequity. A polygenic risk score (PRS) case study illustrates all four domains and shows how a score can misclassify risk in the populations least represented in its derivation while its code, pipeline, and internal validation statistics remain intact. We then translate the framework into practice using ten reviewer prompts that an editor can paste into a review form, a justification template and scoring rubric provided as appendices, a tiered model that triggers full review only for subgroup, equity, transportability, public health, or clinical implementation claims, and a graded account of what should follow an adverse finding. Our argument is that existing governance mechanisms require an upstream layer. Investigators should be asked to justify not only whether their models perform, but whether their representations are truthful enough for the claims at hand. The intended audience is investigators, informaticians, research ethics committees, institutional review boards, data access committees, funders, regulators, and journal editors.
Related Concept Videos
The Representativeness Heuristic
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Bias in Epidemiological Studies
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...