Related Experiment Video
Updated: Jun 4, 2026

Inverse Probability of Treatment Weighting (Propensity Score) using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Using the UMLS and Simple Statistical Methods to Semantically Categorize Causes of Death on Death Certificates
Bill Riedl1, Nhan Than, Michael Hogarth
1Pathology Informatics Core, UC Davis Dept. of Pathology and Laboratory Medicine, UC Davis School of Medicine, Davis, CA.
Abstract:
Cause of death data is an invaluable resource for shaping our understanding of population health. Mortality statistics is one of the principal sources of health information and in many countries the most reliable source of health data. 1 A quick classification process for this data can significantly improve public health efforts. Currently, cause of death data is captured in unstructured form requiring months to process. We think this process can be automated, at least partially, using simple statistical Natural Language Processing, NLP, techniques and the Unified Medical Language System, UMLS, as a vocabulary resource. A system, Medical Match Master, MMM, was built to exercise this theory. We evaluate this simple NLP approach in the classification of causes of death. This technique performed well if we engaged the use of a large biomedical vocabulary and applied certain syntactic maneuvers made possible by textual relationships within the vocabulary.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Life Tables
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...
Applications of Life Tables
Introduction To Survival Analysis
The primary goal of survival analysis is to estimate survival time—the time until a...
Kaplan-Meier Approach
