Related Experiment Video
Updated: Jun 14, 2025

06:18
Optimized Bone Sampling Protocols for the Retrieval of Ancient DNA from Archaeological Remains
Published on: November 30, 2021
3.8K
An open dataset for oracle bone character recognition and decipherment
Pengjie Wang1, Kaile Zhang1, Xinyu Wang2
1Huazhong University of Science and Technology, Wuhan, 430074, China.
Scientific Data
|September 6, 2024
Summary
Researchers created the HUST-OBC dataset to aid in deciphering ancient Chinese Oracle Bone Script (OBC). This large dataset of over 140,000 images aims to overcome challenges in understanding Shang Dynasty texts.
Area of Science:
- Digital Humanities
- Artificial Intelligence
- Archaeology
Background:
- Oracle Bone Script (OBC) is a crucial early Chinese writing system from the Shang Dynasty (3,000 years ago).
- Deciphering OBC is vital for understanding ancient Chinese history and culture, but is challenging due to text degradation.
- Artificial Intelligence (AI) offers potential for OBC decipherment, yet lacks sufficient high-quality datasets.
Purpose of the Study:
- To introduce the HUST-OBC dataset, a comprehensive resource for AI-driven OBC decipherment.
- To facilitate research into ancient Chinese writing and Shang Dynasty studies.
Main Methods:
- Compilation of a large-scale dataset (HUST-OBC) from diverse sources.
- Inclusion of 140,053 images: 77,064 of 1,588 deciphered characters and 62,989 of 9,411 undeciphered characters.
- Making all associated codes and the dataset publicly available.
Main Results:
- Creation of the HUST-OBC dataset, featuring a substantial number of both deciphered and undeciphered Oracle Bone Characters.
- The dataset provides a foundation for developing and training AI models for ancient script analysis.
Conclusions:
- The HUST-OBC dataset addresses the critical need for high-quality data in AI-assisted OBC research.
- This resource is expected to significantly advance the decipherment of ancient Chinese writing and related historical studies.

