A multimodal benchmark dataset for evaluating large language models on traditional Chinese opera understanding
Gengxian Cao1, Shuo Hou2, Ye Yang3
1The Music Education Center at the School of Humanities, Xi'an Jiaotong University, Xi'an, China.
Scientific Data
|June 30, 2026
Summary
Researchers developed the TCO-Dataset, a new bilingual benchmark for evaluating large language models (LLMs) on traditional Chinese opera images. This dataset aids in assessing AI
Area of Science:
- Artificial Intelligence
- Computer Vision
- Cultural Heritage Studies
Background:
- Large language models (LLMs) require robust benchmarking for capability evaluation.
- Existing multimodal benchmarks inadequately cover culturally rich domains like traditional Chinese opera.
- There is a need for specialized datasets to assess AI performance in visual-cultural reasoning.
Purpose of the Study:
- Introduce the TCO-Dataset, a novel bilingual multimodal dataset for traditional Chinese opera.
- Enable the assessment of LLMs' visual-cultural reasoning abilities on complex imagery.
- Facilitate cross-lingual evaluation of AI models in a specialized domain.
Main Methods:
- Curated 1,000 multiple-choice questions with high-resolution images across eight Chinese opera genres.
- Ensured bilingual support (Chinese and English) for diverse applications.
- Implemented multiple rounds of expert validation for data accuracy and consistency.
Main Results:
- The TCO-Dataset presents a challenging benchmark for multimodal AI.
- Initial evaluations reveal significant performance disparities among different LLMs.
- The dataset effectively highlights the need for domain-specific AI development.
Conclusions:
- The TCO-Dataset fills a critical gap in multimodal AI benchmarking for traditional Chinese opera.
- It serves as a valuable resource for advancing AI's understanding of visual-cultural nuances.
- The dataset supports AI development for cultural heritage preservation and domain-specific applications.

