Public preprint

Tracing the Evolution of Oracle Bone Characters Across Three Millennia

We present a continuous-manifold approach to Chinese script evolution, connecting fragmentary oracle-bone evidence, cross-era supervision, Neural ODE transition dynamics, and survival-aware bidirectional decipherment.

Tianhao Fu¹ · Xinxin Xu² · Spike Wang¹ · Cunyi Kang³ · Jian Cao² · Xixin Cao²¹ Fulcrum.AI · ² Peking University · ³ Independent
Abstract

1. Abstract

We introduce MSEF, a framework that treats oracle-bone decipherment as a problem of reconstructing historical script trajectories rather than matching isolated glyph images.

Oracle Bone Inscription decipherment is difficult because many ancient glyphs do not map cleanly to one later script form. Existing computational methods often compare OBI with a single later period independently, but characters may change non-monotonically across Bronze, Seal, Clerical, and modern scripts. We propose the Manifold-based Script Evolution Framework (MSEF), which models Chinese script evolution as continuous movement through era-specific manifolds. A Neural ODE learns inter-era transition dynamics, while a survival-aware bidirectional inference procedure recalls plausible candidates forward and verifies them backward. The associated public CCAMC corpus provides occurrence-level source records with dynasty, period, script, provenance, and image references. It is a source component for this work; the FGCCES alignments, survival annotations, engineered features, and character-disjoint splits are not included in the public archive.

01Continuous evolution. Glyphs are modeled as historical trajectories.
02Fine-grained data. Partial cross-era pairs become supervision.
03Bidirectional inference. CBED recalls forward and verifies backward.
04Mechanistic analysis. Visualizations inspect learned paleographic structure.
Visual intuition

2. Intuition: Script Evolution as a Trajectory

Illustrative trajectories visualize the continuous-evolution modeling assumption; they are not independently verified historical reconstructions.

Continuous script evolution frame
Data foundation

3. Dataset: Turning Fragmentary History into Supervision

The public CCAMC fine-grained corpus contains 158,620 source occurrences across six original script categories. Occurrences are distinct from unique characters, objects, and images. The package includes occurrence tables, cached source pages, and linked source images. It is not the complete FGCCES benchmark: cross-era character alignments, survival labels, engineered feature files, and train/validation/test splits are not part of this release.

The CCAMC records retain the source collection’s script categories, dynasty and period labels, bibliographic context, and image references. They do not encode inferred cross-era correspondences or extinction labels. These source records are distinct from the derived FGCCES annotations used in the paper’s experiments.

  1. 01Source coverage. The archive retains six CCAMC source categories. It does not itself form complete five-era trajectories.
  2. 02Temporal detail. Original subperiod and dynasty labels are retained where present in the source.
  3. 03Source-backed records. No character match is inferred solely from similar names, and missing script forms are not marked extinct.
158,620
released occurrence records
41,931
glyph image files
29,650
cached source pages
Method

4. Method: Diagnose → Model → Verify → Inspect

Each view answers one reader question: why existing approaches fail, how to model evolution, how to verify a candidate, and how to inspect what was learned.

01

Diagnose

Show why direct generation, retrieval, and hand-built rules break when intermediate evidence is missing.

02

Model

Represent each glyph as an era-conditioned manifold state and learn inter-era movement with Neural ODEs.

03

Verify

Recall candidates forward, prune implausible or extinct paths, and check consistency by tracing backward.

04

Inspect

Use mechanistic visualizations to test whether the learned geometry reflects paleographic structure.

Evidence path

5. Read the Figures as the Proof of the Method

The figure groups follow the same logic as the method: diagnostic evidence motivates the model, core diagrams define the framework, inference figures explain verification, and mechanistic figures inspect the learned structure.

1. Pre-Exploratory Diagnosis

Diagnostic evidence

These figures motivate the method: missing stages, branching variants, and rule explosion make isolated matching brittle.

2. MSEF: Continuous Manifold Formulation

Core model

These figures show the learned state space: glyphs become era-conditioned points, and Neural ODEs describe how they move through history.

Supporting method diagramsSupporting evidence

3. CBED: Bidirectional Decipherment

Inference

These figures show the inference path: retrieve broadly first, then remove candidates that are extinct, implausible, or not reconstructable backward.

4. Post-Analysis: What the Model Learns

Mechanistic evidence

These figures test the learned mechanism: where the model attends, what features matter causally, and whether the learned geometry has interpretable structure.

Large visual evidence

6. Dense Figures for Close Inspection

Dense qualitative evidence is presented at larger scale so that neighborhoods, trajectories, retrieval examples, and the global variant network can be inspected without losing local structure.

Results

7. Experimental Results Across Datasets

These tables reproduce the reported manuscript results. They are distinct from the source-data statistics and implementation checks supplied with the code release.

HUST-OBS / EVOBC

Single-round Top-1 accuracy (%).

MethodOBS-OCRPaddleOCR
Pix2Pix00
CycleGAN00
BBDM19.57
CDE3119
OBSD4130
MSEF71.558.5

PictOBI-20k

Four-choice visual association; accuracy (%).

ModelNormalComplexOverall
Random252525
GPT-4o-2024-11-2026.3125.5226.23
Gemini 2.5 Pro55.2239.4453.66
Claude 4 Sonnet35.9325.9234.94
GLM-4.5V-106B33.1927.1132.48
Qwen2.5-VL-72B25.4124.9825.36
InternVL3-78B52.2936.3850.71
InternVL3-38B52.7139.5151.4
MSEF74.8253.2472.18

FGCCES-Test

Mean ± standard deviation across five runs; recall (%).

MethodR@1R@5R@10R@1%AP
Diff-Oracle52.1 ± 1.368.9 ± 1.677.4 ± 1.285.2 ± 1.00.56
OracleFusion58.3 ± 1.274.5 ± 1.381.2 ± 1.188.1 ± 0.90.62
CrossFont54.6 ± 1.470.8 ± 1.478.5 ± 1.286.5 ± 1.00.58
OracleSage60.1 ± 1.176.2 ± 1.282.8 ± 1.089.5 ± 0.80.64
OracleAgent62.8 ± 1.077.9 ± 1.184.6 ± 0.990.8 ± 0.70.67
MSEF72.5 ± 0.886.5 ± 0.691.8 ± 0.595.6 ± 0.40.78
Citation

BibTeX

@misc{fu2026tracing,
  title     = {Tracing the Evolution of Oracle Bone Characters Across Three Millennia},
  author    = {Fu, Tianhao and Xu, Xinxin and Wang, Spike and Kang, Cunyi and Cao, Jian and Cao, Xixin},
  note      = {Preprint},
  year      = {2026}
}
Contact

Get in touch

tianhaofu1@gmail.com