Related Experiment Video
Updated: Jul 6, 2026

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Classifying microfossil radiolarians on fractal pre-trained vision transformers.
Kazuhide Mimura1, Takuya Itaki2,3, Hirokatsu Kataoka4,5
1Geological Survey of Japan, National Institute of Advanced Industrial Science and Technology, 1-1-1 Higashi, Tsukuba, Ibaraki, 305-8567, Japan. kaz-mimura@aist.go.jp.
Deep learning, including vision transformers (ViT) and formula-driven supervised learning (FDSL), shows promise for geological image classification. These methods achieved higher precision in microfossil classification compared to traditional models.
Area of Science:
- Geoscience
- Computer Science
- Artificial Intelligence
Background:
- Deep learning, particularly for image classification, has advanced rapidly but lags in geological applications.
- Convolutional Neural Networks (CNNs) are common, but Vision Transformers (ViT) offer a promising alternative architecture.
- Formula-Driven Supervised Learning (FDSL) uses generated images for pre-training, potentially improving model performance.
Purpose of the Study:
- To evaluate the efficacy of Vision Transformers (ViT) and Formula-Driven Supervised Learning (FDSL) for microfossil image classification.
- To address the time lag in applying advanced deep learning techniques to geological studies.
- To compare the performance of ViT and FDSL models against traditional CNN models in a geological context.
Main Methods:
- Applied Vision Transformer (ViT) architecture for microfossil image classification.
- Utilized Formula-Driven Supervised Learning (FDSL) for pre-training classification models with generated images.
- Compared ViT and FDSL performance against a baseline Convolutional Neural Network (CNN) model.
Main Results:
- The ViT-based model demonstrated a 6-8% higher average precision compared to the previous CNN model.
- Models pre-trained using FDSL exhibited slightly higher precision on average than those pre-trained on real images.
- ViT and FDSL techniques proved effective for classifying microfossils (radiolarians).
Conclusions:
- Vision Transformers (ViT) and Formula-Driven Supervised Learning (FDSL) are suitable for geological image classification tasks.
- These advanced deep learning techniques can improve the accuracy and efficiency of analyzing geological imagery.
- The study highlights the potential of ViT and FDSL to accelerate the adoption of deep learning in geosciences.
Related Concept Videos
Transformers with Off-Nominal Turns Ratios
Diversity of Protists III

