Related Experiment Video
Updated: Sep 14, 2025

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Autoencoder-Based Representation Learning for Similar Patients Retrieval From Electronic Health Records: Comparative
Deyi Li1, Aditi Shukla2, Sravani Chandaka3
1Department of Health Outcomes & Biomedical Informatics, University of Florida, 1889 Museum Rd, 7th Floor, Suite 7000, Room 7012, Gainesville, FL, 32611, United States, 1 352-627-9143.
Denoising autoencoders excel at finding similar patients using Euclidean distance, outperforming other models. Learning rates and distance measures significantly impact performance in patient representation learning for personalized medicine.
Area of Science:
- Medical informatics
- Machine learning in healthcare
- Computational biology
Background:
- Electronic health record (EHR) data modeling is challenging due to high dimensionality, mixed features, noise, bias, and sparsity.
- Patient representation learning using autoencoders (AEs) offers a promising approach to address these EHR data challenges.
- Understanding the impact of different AE designs and distance measures on retrieving similar patient cohorts is crucial.
Purpose of the Study:
- To evaluate the performance of five common autoencoder (AE) variants in retrieving similar patients.
- To investigate the influence of various distance measures and hyperparameter configurations on AE model performance.
- To assess the effectiveness of AE-based patient similarity estimation for clinical outcome prediction.
Main Methods:
- Tested five AE variants (vanilla, denoising, contractive, sparse, robust) on two real-world EHR datasets.
- Applied k-nearest neighbors (k-NN) with Euclidean and Mahalanobis distances to AE-produced latent representations.
- Evaluated model performance on predicting acute kidney injury onset and 1-year postdischarge mortality.
Main Results:
- Denoising autoencoders significantly outperformed other AE variants when paired with Euclidean distance (P<.001).
- Learning rates were identified as a critical hyperparameter influencing AE model performance.
- Mahalanobis distance-based k-NN frequently outperformed Euclidean distance-based k-NN on latent representations, though direct application to raw data was data-dependent.
Conclusions:
- This study provides a comprehensive analysis of AE variants for patient similarity retrieval.
- Findings highlight the importance of AE design and hyperparameter tuning for effective patient representation learning.
- The results lay the groundwork for developing advanced AE-based methods for personalized medicine and improved patient care.
More Related Videos
09:00TBase - an Integrated Electronic Health Record and Research Database for Kidney Transplant Recipients
Published on: April 13, 2021
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Methods of Documentation VII: EMR
Patient-centered Care
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Electrocardiogram
Three major waveforms are present in a typical ECG recording: the P wave, the QRS complex, and...
Purpose of Health Records II
Ethical Standards II
Nurses are entrusted with upholding various ethical principles and standards. Nurses forge solid therapeutic relationships using trust, empathy, autonomy, confidentiality, and professional competence.
Confidentiality is crucial, embodying respect for individual privacy...