Related Experiment Video
Updated: Jan 8, 2026

Mouse Footpad Inoculation Model to Study Viral-Induced Neuroinflammatory Responses
Published on: June 14, 2020
Basic Science and Pathogenesis
Sungjoon Park1, Kyungwook Lee1, Soorin Yim1
1LG AI Research, Gangseo-gu, Seoul, Korea, Republic of (South).
Background:
Multi-omics data from large-scale consortium databases, such as the Religious Orders Study and Rush Memory and Aging Project (ROSMAP), combined with advanced AI technologies, hold significant promise for identifying biological mechanisms and biomarkers for diagnosis and treatment. However, two major challenges hinder effective multi-omics integration: modality collapse, where prediction models overly rely on a single modality, and data incompleteness, which limits the potential of machine learning approaches to fully utilize these rich resources. To address these issues, we developed an Alzheimer's disease (AD) prediction method capable of utilizing incomplete modalities while identifying key biomarkers through feature importance analysis.
Method:
The proposed model is designed to ensure that each modality contributes meaningfully to phenotype prediction, even in the absence of some modalities. It comprises three modules: Encoder, Aggregator, and Predictor. Each omics data type is independently encoded into an embedding vector, which is then aggregated into a unified representation to facilitate the integration of diverse data types for AD prediction. This unified vector, along with other modality-specific embedding vectors, is fed to a shared predictor. To align the heterogeneous omics embeddings, the model computes a collective loss that integrates both the unified and modality-specific vectors, akin to main and auxiliary tasks in multi-task learning. For missing modalities, the embedding vectors of other available modalities are amplified to compensate for the loss of information.
Result:
Trained on multi-omics samples with both complete and incomplete modalities, the model achieved an accuracy of 0.890 in classifying CogDX labels, outperforming existing state-of-the-art models. Ablation studies showed that all omics data contributed to the prediction, effectively preventing modality collapse. Including samples with missing modalities significantly boosted performance, emphasizing the importance of leveraging incomplete data. The model's performance was also evaluated across varying numbers of available modalities to present its practicality. Feature importance analysis identified biomarkers consistent with findings in existing literature.
Conclusion:
The proposed method effectively integrates multi-omics data, even with incomplete modalities. By rediscovering biomarkers aligned with other studies, it demonstrates the potential of deep learning approaches in multi-omics research for AD.
Related Concept Videos
Infection
The chain begins with pathogens: bacteria, viruses, fungi, prions, or parasites such as protozoa helminths. These can be present on the skin as transient or resident flora, or they can be acquired from the environment. Identifying and treating the type of infection and...
Urinary Tract Infection II: Pathophysiology
Cystic Fibrosis: Pathogenesis
CF is primarily caused by a genetic mutation in a chromosome 7 gene coding for the cystic fibrosis transmembrane conductance regulator (CFTR) protein. The most common gene mutation leading to CF is the ΔF508 mutation,...
Pneumonia II: Pathophysiology
Stages of Infection
Defense Against Bacterial Pathogens
Phagocytes
Phagocytes are the frontline soldiers of the immune system. They include neutrophils and macrophages. Neutrophils are the most abundant type of white blood cell and are quickly mobilized to the site of infection. Macrophages are larger cells that patrol...

