CLWD: a Chinese histopathology dataset for lung adenocarcinoma subtype classification
Yang Chen1, Haoyun Zhao2, Li Wang1
1Department of Pathology, The First People's Hospital of Yunnan Province/The Affiliated Hospital of Kunming University of Science and Technology, Kunming, 650032, Yunnan, China.
Scientific Data
|March 5, 2026
Summary
The CLWD dataset offers 408 whole-slide images for lung adenocarcinoma subtype research. This valuable resource aids in improving diagnostic accuracy for lung cancer patients globally.
Area of Science:
- Pathology
- Oncology
- Medical Imaging
Background:
- Accurate typing, subtyping, and grading are crucial for effective lung adenocarcinoma diagnosis and treatment.
- Existing datasets may not fully represent diverse patient demographics, particularly in Asian populations.
Purpose of the Study:
- Introduce the CLWD dataset, a large-scale resource for lung adenocarcinoma subtype research.
- Facilitate machine learning model development and validation for improved diagnostic accuracy.
- Support research focusing on Chinese patient demographics.
Main Methods:
- Curated a dataset of 408 whole-slide images (WSIs) from 210 lung adenocarcinoma patients.
- Scanned WSIs at 80× magnification.
- Included comprehensive clinical information (age, sex, diagnosis) for each patient.
Main Results:
- The CLWD dataset is one of the largest Asian datasets for lung adenocarcinoma research.
- Initial evaluation using a multi-instance learning framework showed promise for subtype classification.
- The dataset's public accessibility supports diverse research applications.
Conclusions:
- The CLWD dataset is a significant resource for advancing lung cancer pathology research.
- It has the potential to improve the accuracy of lung adenocarcinoma subtype diagnosis globally.
- Facilitates machine learning applications in cancer diagnostics.


