Related Experiment Video
Updated: Jul 10, 2026

06:52
Optical Clearing and Labeling for Light-sheet Fluorescence Microscopy in Large-scale Human Brain Imaging
Published on: January 26, 2024
A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning
Stefano Cerri1,2,3, Asbjørn Munk4,5, Sebastian Nørgaard Llambias6,7
1Department of Computer Science, University of Copenhagen, Copenhagen, Denmark. stce@di.ku.dk.
Scientific Data
|July 8, 2026
Summary
We introduce FOMO260K, a large dataset of 260,927 brain Magnetic Resonance Imaging (MRI) scans. This resource supports developing self-supervised learning methods for medical imaging analysis.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Neuroscience
Background:
- Large-scale datasets are crucial for advancing deep learning in medical imaging.
- Existing datasets often lack diversity in image types and anatomical variability.
- Developing robust self-supervised learning (SSL) models requires extensive and heterogeneous data.
Purpose of the Study:
- To introduce FOMO260K, a comprehensive dataset for medical imaging research.
- To facilitate the development and benchmarking of SSL methods in brain MRI analysis.
- To provide a large-scale, heterogeneous resource with minimal preprocessing.
Main Methods:
- Aggregated 260,927 brain MRI scans from 910 public sources.
- Included clinical- and research-grade images across multiple MRI sequences.
- Applied minimal preprocessing to preserve original image characteristics.
Main Results:
- Created FOMO260K, comprising scans from 55,378 subjects.
- Dataset exhibits wide anatomical and pathological variability, including brain anomalies.
- Provided companion code and pretrained models for SSL tasks.
Conclusions:
- FOMO260K enables large-scale development and benchmarking of SSL in medical imaging.
- The dataset's heterogeneity supports robust model training.
- Availability of code and models accelerates research in AI for neuroimaging.
