Related Experiment Video
Updated: Jan 9, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
A dataset for health insurance analysis: Integrating individual and area-based contextual variables
Josep Lledó1, Priscila Espinosa2, Virgilio Pérez2
1Department of Applied Economics, University of Valencia, Avda. Tarongers s/n, Valencia, 46022, Spain. Josep.Lledo@uv.es.
None:
Access to real data is a challenge for research and professional analysis in the insurance sector, especially since such access is often restricted due to confidentiality and competitive issues, particularly in health insurance. This paper introduces a new dataset from a Spanish health insurance portfolio, covering the years 2017 to 2019 with over 70 thousand unique insured and more than 225 thousand rows of data. The dataset contains 42 variables, of which 27 are directly sourced from the insurer and the rest are derived to include area-based contextual information obtained from publicly available sources. The data are anonymized to ensure privacy while maintaining the integrity required for robust professional analysis. Researchers can use the dataset to explore health insurance dynamics, from product design to contextual effects and risk management. Moreover, it supports academic applications for students and educators where they can use real-world data for exercises in data cleaning, statistical analysis and machine learning models.
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Dimensions of Health and Illness
Biostatistics: Overview
Discrete variables are...
Statistical Software for Data Analysis and Clinical Trials
Study Designs in Epidemiology
Observational studies are those where the researcher does not intervene but rather observes natural variations. They include cross-sectional, cohort, and...
Bias in Epidemiological Studies

