With the support of big data and AI, oracle bone script research has entered a new era. This repository curates open datasets, benchmarks, codebases, and papers for AI-assisted oracle bone inscription recognition, retrieval, rejoining, decipherment, and interpretation.
Last maintained: 2026-07-02
Comprehensive paper index
- Recent Oracle Bone Projects and Papers
- Datasets and Benchmarks
- Paper Index
- Online Resources
- Our Projects
- Contributing
- Copyright
This section lists representative recent oracle bone inscription work from all groups, including our own projects, ordered by year. See PAPERS.md for the more detailed task-oriented paper index.
| Year | Project or Paper | Venue and Status | Category | Links |
|---|---|---|---|---|
| 2026 | AlphaOracle | The Innovation | Decipherment and interpretation | Paper, Code |
| 2026 | Oracle Bone Inscriptions Information Processing: A Comprehensive Survey | npj Heritage Science 2026 | Survey and resources | Paper, Repo |
| 2026 | OBIMD: A Multi-modal Dataset for Contextual Interpretation of Oracle Bone Inscriptions | Scientific Data 2026 | Multimodal dataset | Paper, arXiv, Code, HF |
| 2026 | PictOBI-20k | ICASSP 2026 | Visual decipherment benchmark | IEEE, arXiv, Code |
| 2026 | Chronicles-OCR | arXiv 2026 | Cross-temporal OCR benchmark | arXiv, Code, HF |
| 2026 | Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding | arXiv 2026 | Sentence-level OBI benchmark | arXiv |
| 2026 | Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention | arXiv 2026 | Recognition | arXiv |
| 2026 | OracleAnalyser | arXiv 2026 | MLLM-based oracle-bone analysis | arXiv |
| 2026 | Oracle bone inscription detection model with frequency-domain attention fusion and multi-scale optimization | npj Heritage Science 2026 | Detection | Paper |
| 2026 | Explainable Oracle Bone Script Recognition via Multimodal Pictographic Reasoning | AAAI 2026 | Explainable recognition | Paper |
| 2026 | OracleDet | npj Heritage Science 2026 | Complex-scene OBI detection | Paper, Code |
| 2026 | OBI Designer | npj Heritage Science 2026 | Artistic OBI character generation | Paper |
| 2026 | Specializing Large Models for OBS Interpretation via Component-Grounded Multimodal Knowledge Augmentation | arXiv 2026 | Knowledge-augmented interpretation | arXiv |
| 2026 | Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval | arXiv 2026 | Dictionary retrieval | arXiv |
| 2025 | OracleFusion | ICCV 2025 | Structurally constrained semantic typography | Paper, arXiv, Code |
| 2025 | V-Oracle | ACL 2025 | Progressive VQA-style reasoning | Paper |
| 2025 | OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones? | ICLR 2025 Spotlight | LMM benchmark | OpenReview, arXiv, Code |
| 2025 | OracleAgent | arXiv 2025 | Multimodal research agent | Paper, Code |
| 2025 | Interpretable OBS Decipherment with LVLMs, PD-OBS | arXiv 2025 | Interpretable decipherment | Paper, Code |
| 2025 | Oracle-P15K | ACM MM 2025 | Long-tail recognition dataset and benchmark | Paper, Code |
| 2025 | OBIFormer | Displays 2025 | Denoising and restoration | Paper, Code |
| 2025 | A Graph-based Evolutionary Dataset for Oracle Bone Characters | npj Heritage Science 2025 | Character evolution graph | Paper, Code |
| 2025 | An Open Benchmark for Oracle Bone Rubbing Image Retrieval | npj Heritage Science 2025 | Rubbing-image retrieval | Paper |
| 2025 | Deep Rejoining Model and Dataset of Oracle Bone Fragment Images | npj Heritage Science 2025 | Fragment rejoining | Paper |
| 2025 | A Text-Image Dual Conditional Stable Diffusion Model for OBI Decipherment | npj Heritage Science 2025 | Dual conditional diffusion | Paper |
| 2025 | A Cross-Font Image Retrieval Network for Recognizing Undeciphered OBI | ICIC 2025, arXiv 2024 | Cross-font retrieval | arXiv |
| 2025 | Component-Level Segmentation for OBI Decipherment | AAAI 2025 | Component segmentation | Paper, Code |
| 2024 | OBSD: Deciphering Oracle Bone Language with Diffusion Models | ACL 2024 Best Paper | Diffusion-based decipherment | Paper, arXiv, Code |
| 2024 | Puzzle Pieces Picker (P3) | ICDAR 2024 Oral | Radical and stroke reconstruction | Paper, Code |
| 2024 | HUST-OBC | Scientific Data 2024 | Recognition and decipherment dataset | Paper, arXiv, Code, Data |
| 2024 | EVOBC | arXiv 2024 | Multi-period character evolution dataset | Paper, Code, Data |
| 2024 | OracleSage | arXiv 2024 | Visual-linguistic understanding | Paper |
| Dataset or Benchmark | Task | Scale and Notes | Links |
|---|---|---|---|
| HUST-OBC | Recognition + decipherment | 140,053 images; deciphered and undeciphered character categories | Paper, Code, Data |
| EVOBC | Character evolution | Multi-period evolution data across OBC, BI, SS, SAC, WSC, CS | Paper, Code, Data |
| OBIMD | Contextual interpretation | 10,077 OBI images; 93,652 annotated characters; sentence-level readings | Paper, Code, HF |
| OBI-Bench | LMM benchmark | Five OBI processing tasks; 5,523 images | OpenReview, Code |
| PictOBI-20k | Visual decipherment | 20k OBC-object image pairs; 15k+ multi-choice questions | IEEE, arXiv, Code |
| Chronicles-OCR | Cross-temporal OCR | 2,800 balanced images across the Seven Chinese Scripts, including oracle bone script | arXiv, Code, HF |
| S-OBI | Sentence-level OBI understanding | 95 standardized sentence-level OBI instances and 695 QA pairs for semantic matching, slot extraction, and contextual reasoning | arXiv |
| Oracle-MNIST | Benchmark classification | 30,222 grayscale oracle-character images in 10 categories | Paper, Code |
| Oracle-P15K | Long-tail recognition | Long-tail OBI benchmark with synthesis-based augmentation | Paper, Code |
| PD-OBS | Interpretable decipherment | Radical and pictographic annotations for LVLM training | Paper, Code |
| GEVOBC and GBEDOBC | Evolution graph dataset | Graph-based evolutionary oracle bone character dataset | Paper, Code |
| OBI Rubbing Retrieval Benchmark | Rubbing retrieval | Homologous rubbing retrieval benchmark | Paper |
| OBFI and OBID-ACR | Fragment rejoining and bone association | Fragment-image and bone-level association benchmarks | OBFI, OBID-ACR |
| OBC306, HWOBC, YinQiWenYuan | Recognition and detection | Public resources from Yin Qi Wen Yuan | Website |
This README highlights representative works. For a broader task-oriented bibliography, see:
- ๐
PAPERS.md: surveys, knowledge resources, datasets, decipherment, multimodal reasoning, recognition, detection, segmentation, retrieval, rejoining, restoration, generation, and broader ancient-script processing. - ๐ Recommended companion survey repo: OBI-Survey.
| Resource | Link |
|---|---|
| ๆฎทๅฅๆๆธ (Yin Qi Wen Yuan) | Website |
| ๅฐๅญฆๅ (Xiao Xue Tang) | Website |
| ๅฝๅญฆๅคงๅธ (Guo Xue Da Shi) | Website |
| ็ผ็่็ (Zhui Yu Lian Zhu) | Website |
| Yin Xu OBI Database | Database |
| OBI AI Collaborative Platform | Website |
| Multi-function Chinese Character Database | Database |
| Chinese Etymology | Website |
| Omniglot: Oracle Bone Script | Website |
| Museum | Link |
|---|---|
| ๆ ๅฎซๅ็ฉ้ข | Collection |
| ๆฒณๅๅ็ฉ้ข | Collection |
| ่พฝๅฎ็ๅ็ฉ้ฆ | Collection |
| ๅฑฑไธๅ็ฉ้ฆ | Collection |
| ้่ฅฟๅๅฒๅ็ฉ้ฆ | Collection |
| ไธๆตทๅ็ฉ้ฆ | Collection |
| ๆฎทๅขๅ็ฉ้ฆ | Collection |
| ๆตๆฑ็ๅ็ฉ้ฆ | Collection |
| ไธญๅฝๅฝๅฎถๅ็ฉ้ฆ | Collection |
| ้ๅบไธญๅฝไธๅณกๅ็ฉ้ฆ | Collection |
๐ [The Innovation] AlphaOracle: Oracle Bone Script Decipherment via Human-Workflow-Inspired Deep Learning
Yuliang Liu, Haisu Guan, Pengjie Wang, Xinyu Wang, Jinpeng Wan, Kaile Zhang, Handong Zheng, Xingchen Liu, Zhebin Kuang, Huanxin Yang, Bang Li, Yongge Liu, Lianwen Jin, Xiang Bai.
AlphaOracle is published in The Innovation. AlphaOracle integrates computer vision, computational linguistics, and philological validation into a human-workflow-inspired framework for oracle bone script analysis and decipherment.
@article{liu2026alphaoracle,
title = {AlphaOracle: Oracle Bone Script Decipherment via Human-Workflow-Inspired Deep Learning},
author = {Liu, Yuliang and Guan, Haisu and Wang, Pengjie and Wang, Xinyu and Wan, Jinpeng and Zhang, Kaile and Zheng, Handong and Liu, Xingchen and Kuang, Zhebin and Yang, Huanxin and Li, Bang and Liu, Yongge and Jin, Lianwen and Bai, Xiang},
journal = {The Innovation},
year = {2026},
url = {https://www.sciencedirect.com/science/article/pii/S2666675826002092}
}Haisu Guan, Huanxin Yang, Xinyu Wang, Shengwei Han, Yongge Liu, Lianwen Jin, Xiang Bai, Yuliang Liu.
Paper | arXiv | GitHub Code
This paper introduces Oracle Bone Script Decipher (OBSD), a conditional diffusion-based strategy that generates modern-character clues for oracle bone script decipherment.
@inproceedings{guan2024deciphering,
title = {Deciphering Oracle Bone Language with Diffusion Models},
author = {Guan, Haisu and Yang, Huanxin and Wang, Xinyu and Han, Shengwei and Liu, Yongge and Jin, Lianwen and Bai, Xiang and Liu, Yuliang},
booktitle = {Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
pages = {15554--15567},
year = {2024},
doi = {10.18653/v1/2024.acl-long.831}
}๐ [ICDAR 2024 Oral] Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction
Pengjie Wang, Kaile Zhang, Xinyu Wang, Shengwei Han, Yongge Liu, Lianwen Jin, Xiang Bai, Yuliang Liu.
Puzzle Pieces Picker (P3) deconstructs oracle bone inscriptions into strokes and radicals and reconstructs them into modern counterparts with a Transformer-based model.
@inproceedings{wang2024puzzle,
title = {Puzzle Pieces Picker: Deciphering Ancient Chinese Characters with Radical Reconstruction},
author = {Wang, Pengjie and Zhang, Kaile and Wang, Xinyu and Han, Shengwei and Liu, Yongge and Jin, Lianwen and Bai, Xiang and Liu, Yuliang},
booktitle = {Document Analysis and Recognition -- ICDAR 2024},
year = {2024},
publisher = {Springer},
doi = {10.1007/978-3-031-70533-5_11}
}๐ [SCIENTIA SINICA Informationis 2026] A Multi-task Multimodal Reasoning Framework for Oracle Bone Character Interpretation
Jinpeng Wan, Yuliang Liu, Xiang Bai.
Paper | GitHub Code | Data
This paper introduces ORACLE-PRIME, a multi-task multimodal reasoning framework for oracle bone character interpretation. It integrates glyph perception, structural periodization, and evolutionary reasoning to generate philologically grounded reasoning chains.
@article{wan2026oracleprime,
title = {A multi-task multimodal reasoning framework for Oracle Bone character interpretation},
author = {Wan, Jinpeng and Liu, Yuliang and Bai, Xiang},
journal = {SCIENTIA SINICA Informationis},
year = {2026},
pages = {-},
url = {http://www.sciengine.com/publisher/Science China Press/journal/SCIENTIA SINICA Informationis///10.1360/SSI-2025-0551},
doi = {10.1360/SSI-2025-0551}
}Haisu Guan, Jinpeng Wan, Yuliang Liu, Pengjie Wang, Kaile Zhang, Zhebin Kuang, Xinyu Wang, Xiang Bai, Lianwen Jin.
Paper | GitHub Code | Data
EVOBC collects character images across six historical stages: Oracle Bone Characters, Bronze Inscriptions, Seal Script, Spring and Autumn period characters, Warring States period characters, and Clerical Script.
@article{guan2024open,
title = {An Open Dataset for the Evolution of Oracle Bone Characters: EVOBC},
author = {Guan, Haisu and Wan, Jinpeng and Liu, Yuliang and Wang, Pengjie and Zhang, Kaile and Kuang, Zhebin and Wang, Xinyu and Bai, Xiang and Jin, Lianwen},
journal = {arXiv preprint arXiv:2401.12467},
year = {2024}
}Pengjie Wang, Kaile Zhang, Xinyu Wang, Shengwei Han, Yongge Liu, Jinpeng Wan, Haisu Guan, Zhebin Kuang, Lianwen Jin, Xiang Bai, Yuliang Liu.
Paper | arXiv | GitHub Code | Data | hyper.ai
HUST-OBC is a large-scale open dataset for oracle bone character recognition and decipherment, containing deciphered and undeciphered OBC images.
@article{wang2024open,
title = {An Open Dataset for Oracle Bone Character Recognition and Decipherment},
author = {Wang, Pengjie and Zhang, Kaile and Wang, Xinyu and Han, Shengwei and Liu, Yongge and Wan, Jinpeng and Guan, Haisu and Kuang, Zhebin and Jin, Lianwen and Bai, Xiang and Liu, Yuliang},
journal = {Scientific Data},
volume = {11},
number = {1},
year = {2024},
doi = {10.1038/s41597-024-03807-x}
}We welcome pull requests and issues. For new papers and resources, please include:
- Title, authors, venue and year, and task category.
- Official paper link, DOI, arXiv, or OpenReview page when available.
- Code, data, or demo links if public.
- A one-line summary of the contribution.
We welcome suggestions to help improve Open-Oracle. For any query, please contact Prof. Yuliang Liu: ylliu@hust.edu.cn. If you find something interesting, feel free to share it by email or open an issue. Thanks!






