Medical multimodal multitask foundation model for lung cancer screening
Chuang Niu1, Qing Lyu2, Christopher D Carothers1
1Department of Biomedical Engineering, School of Engineering, Biomedical Imaging Center, Center for Computational Innovations, Center for Biotechnology & Interdisciplinary Studies, Rensselaer Polytechnic Institute, 110 8th Street, Troy, 12180, NY, USA.
Nature Communications
|February 11, 2025
Summary
A new medical foundation model (M3FM) enhances lung cancer screening by analyzing diverse data types. This AI approach improves lung cancer and cardiovascular disease risk prediction, advancing clinical management.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Multimodal Data Fusion
Background:
- Lung cancer screening (LCS) utilizes extensive multimodal data, including text, tables, and images.
- Incomplete data mining in LCS can lead to overlooked features, negatively impacting patient care and outcomes.
- The complexity of LCS data necessitates advanced analytical approaches to maximize clinical utility.
Purpose of the Study:
- To introduce a novel medical multimodal-multitask foundation model (M3FM) for analyzing three-dimensional low-dose computed tomography (CT) data in LCS.
- To develop a scalable architecture capable of synergistic multimodal multitasking for comprehensive LCS data analysis.
- To improve the accuracy and efficiency of various LCS tasks through advanced AI techniques.
Main Methods:
- Curated a large-scale dataset comprising 49 clinical data types, 163,725 chest CT series, and 17 distinct LCS tasks.
- Developed a scalable multimodal question-answering model architecture designed for synergistic multitasking.
- Employed large-scale multimodal and multitask learning strategies to train the M3FM.
Main Results:
- M3FM demonstrated superior performance compared to existing state-of-the-art models in LCS.
- Achieved significant improvements in lung cancer risk prediction (up to 20%) and cardiovascular disease mortality risk prediction (up to 10%).
- The model effectively processes high-dimensional, multiscale imaging data and integrates diverse data modalities.
Conclusions:
- M3FM advances LCS by effectively leveraging large-scale multimodal and multitask learning.
- The model exhibits adaptability to out-of-distribution tasks with minimal data requirements.
- This foundation model holds potential for enhancing clinical management and healthcare quality in lung cancer screening.


