Related Experiment Video
Updated: Jun 29, 2025

Using Cholesky Decomposition to Explore Individual Differences in Longitudinal Relations between Reading Skills
Published on: September 17, 2019
Variable selection for latent class analysis in the presence of missing data with application to record linkage
Huiping Xu1, Xiaochun Li1, Zuoyi Zhang2
1Department of Biostatistics and Health Data Science, Indiana University, Indianapolis, IN, USA.
This study introduces a new method for selecting relevant fields in probabilistic record linkage, improving accuracy. The approach effectively handles missing data, leading to better matching performance in real-world applications.
Area of Science:
- Statistics
- Data Science
- Computer Science
Background:
- The Fellegi-Sunter model is a standard for probabilistic record linkage, identifying duplicate records.
- Practitioners often use all available fields, assuming more data improves matching accuracy.
- However, including irrelevant or noisy variables can degrade performance in model-based clustering and linkage.
Purpose of the Study:
- To develop a variable selection procedure for probabilistic record linkage that accounts for missing data.
- To improve the performance of record linkage algorithms by identifying the most informative matching fields.
Main Methods:
- Modified the stepwise variable selection procedure by Fop, Smart, and Murphy.
- Extended the procedure to handle missing data, a common issue in record linkage.
- Evaluated the method using simulation studies and a real-world application.
Main Results:
- The proposed method successfully selected the correct set of matching fields across various scenarios.
- Algorithms using the selected fields demonstrated improved match performance compared to traditional methods.
- The effectiveness was validated in a practical, real-world record linkage scenario.
Conclusions:
- Variable selection is crucial for optimizing probabilistic record linkage, especially when dealing with missing data.
- The proposed modified stepwise procedure effectively identifies informative fields, enhancing linkage accuracy.
- This method is recommended for practitioners seeking to improve their probabilistic record linkage algorithms.
More Related Videos
Related Concept Videos
Friedman Two-way Analysis of Variance by Ranks
Law of Independent Assortment
Comparing the Survival Analysis of Two or More Groups
Contingency Table
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Longitudinal Studies

