Related Experiment Video
Updated: May 4, 2026

Computer-Generated Animal Model Stimuli
Published on: July 29, 2007
A Museum artifact classification model based on cross-modal attention fusion and generative data augmentation.
1School of Culture and Museology, Sichuan Vocational College of Cultural Industries, Chengdu, 610213, China.
This study introduces a novel museum artifact classification model using cross-modal attention and generative data augmentation to improve accuracy and efficiency. The VBG Model enhances cultural heritage preservation by addressing data scarcity and multimodal challenges.
Area of Science:
- Computer Science
- Artificial Intelligence
- Digital Humanities
Background:
- Museum artifact classification is crucial for cultural heritage preservation but is hindered by limited multimodal data and annotation scarcity.
- Traditional and single-modality deep learning models lack the efficiency and accuracy needed for complex artifact classification tasks.
- Integrating diverse data sources and augmenting datasets are key challenges in developing robust classification systems.
Purpose of the Study:
- To propose a novel museum artifact classification model (VBG Model) that overcomes limitations of existing methods.
- To enhance classification accuracy and efficiency by leveraging multimodal information and generative data augmentation.
- To provide a technical solution for digital artifact management and the broader field of cultural heritage preservation.
Main Methods:
- Developed a multimodal framework by integrating Vision Transformer (ViT) for visual features and BERT for textual semantics.
- Implemented a bidirectional interactive attention fusion layer for precise feature alignment between modalities.
- Utilized a Generative Adversarial Network (GAN) for data augmentation, creating a 'generation-feedback-optimization' loop to address data scarcity.
Main Results:
- The VBG Model achieved high performance on MET and MS COCO datasets, with classification accuracies of 92% and 90%, respectively.
- Achieved competitive mAP (0.85 and 0.83) and F1 scores (88% and 86%), outperforming established models like ResNet and DenseNet.
- Ablation studies confirmed the critical contribution of cross-modal fusion and generative data augmentation, with accuracy drops of 5%-9% upon their removal.
Conclusions:
- The proposed VBG Model effectively addresses multimodal data challenges and data scarcity in museum artifact classification.
- Cross-modal attention fusion and generative data augmentation are vital components for achieving high performance.
- Future work will focus on optimizing training time and generated image quality for enhanced artifact distinction and digital preservation.
Related Concept Videos
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Classification of Bones
Long and Short Bones
The appendicular skeleton, particularly the upper and lower limbs, is primarily made of long and short bones. The...
Classification of Systems-II
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Methods of Classification and Identification

