Related Experiment Video
Updated: Aug 13, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.9K
MedMNIST v2 - A large-scale lightweight benchmark for 2D and 3D biomedical image classification
Jiancheng Yang1, Rui Shi1, Donglai Wei2
1Shanghai Jiao Tong University, Shanghai, China.
Scientific Data
|January 19, 2023
Summary
MedMNIST v2 is a large dataset of 2D and 3D biomedical images for machine learning. This collection enables research in biomedical image analysis and computer vision without requiring prior domain expertise.
Area of Science:
- Medical Image Analysis
- Computer Vision
- Machine Learning
Background:
- Standardized biomedical datasets are crucial for developing and validating machine learning models.
- Existing datasets often lack standardization or require specialized domain knowledge.
- A need exists for a comprehensive, accessible collection of biomedical images for diverse research applications.
Purpose of the Study:
- Introduce MedMNIST v2, a large-scale, standardized dataset of 2D and 3D biomedical images.
- Provide a versatile resource for classification tasks across various dataset scales and complexities.
- Facilitate research and education in biomedical image analysis, computer vision, and machine learning.
Main Methods:
- Compiled 18 datasets (12 2D, 6 3D) of biomedical images.
- Pre-processed all images to a uniform size (28x28 for 2D, 28x28x28 for 3D).
- Included corresponding classification labels for each image, ensuring no background knowledge is needed.
Main Results:
- MedMNIST v2 comprises 708,069 2D images and 9,998 3D images.
- The dataset supports diverse tasks: binary/multi-class classification, ordinal regression, and multi-label classification.
- Benchmarking of baseline methods using 2D/3D neural networks and AutoML tools was performed.
Conclusions:
- MedMNIST v2 offers a valuable, accessible resource for advancing biomedical image analysis and machine learning.
- The standardized nature and diverse tasks make it suitable for a wide range of research and educational purposes.
- Public availability of data and code promotes collaboration and innovation in the field.

