Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions data/corpora/guanyin/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
__pycache__/
*.pyc
101 changes: 101 additions & 0 deletions data/corpora/guanyin/INTERPRETATION-LAYER.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
# 觀音百首 — Historical Interpretation Layer

**corpus:** guanyin | **layer:** historical_interpretation | **edition:** 艋舺龍山寺《觀世音靈籤》(薛皓文 2008 附錄二所收錄之籤紙影像)

## 收錄範圍

- 同版兩個欄位:**解曰**(解,四言判詞)+ **聖意**(事項判詞,家宅/自身/求財/交易/婚姻/六甲/行人/田蠶/六畜/尋人/公訟/移徙/失物/疾病/山墳等)
- 只逐字收附錄二既有內容,不收其他網站解釋、不做現代白話、不摘要、不改寫、不補 AI 解讀

## Source hierarchy(production witness 鐵律)

1. **薛皓文 2008 附錄二原始影像 = production witness(唯一權威)**
2. 附錄二的兩個 OCR(整頁「此籤」行 + 逐籤渲染圖 OCR)是 transcription **候選**,不是答案——直排掃描有亂序+形近字誤讀
3. **chance.org.tw 只是 comparison witness**:只協助發現異文、檢查事項項目與順序,**不拿 chance 覆蓋或反推附錄二**

## Coverage

| 項目 | 數量 |
|---|---|
| 籤數 | 100/100 |
| 解曰 | 100 筆 |
| 聖意 | 100 筆 |
| 總 entry | 200 筆 |
| 缺欄位 | 0 |

## 資料結構

每筆 entry 含:`corpus` / `slip_no` / `edition` / `field_type`(解曰|聖意)/ `verbatim_text` / `source_locator` / `transcription_status` / `layer_class`(=living_tradition)/ `variants_or_notes`

**transcription_status 分布:解曰 79 PROBABLE + 21 UNRESOLVED;聖意 100 PROBABLE(verbatim 為附錄二 raw)**

## Authority Hierarchy 窄修記錄(2026-08 Gate FIX REQUIRED 後)

福 Gate 點名兩類問題,已逐一修正:

1. **chance/語義推定曾進入 verbatim**(違反 authority hierarchy)——#1、#5、#7、#13、#16、#18、#25、#41、#48 共 9 籤,將 chance 讀法/語義補字移出 verbatim,改以附錄二 raw/page 為準,chance 只留 witness。
2. **UNRESOLVED 未保留 uncertainty**——#9、#10、#12、#14、#17、#20、#21、#44、#65、#96 共 10 籤,verbatim 改為保留 raw 原讀+缺字「□」,不再以語義推定補滿。

另:#4 第二句「騎龍跨虎」raw「時龍跨虎」/page「騎龍跨虎」/chance「騎龍踏虎」/source image「財施跨虎」四方分歧,標「□龍跨虎」不視為穩定。

**修正後 5 籤由 PROBABLE 降 UNRESOLVED**(#1、#4、#7、#13、#16)。

## Data Gate(機器可抓 contract)

`validate_interpretation_layer.py` 六項 assertion,把 authority hierarchy 變成可回歸驗證的 gate:

| Assertion | 檢查內容 |
|---|---|
| A1 character coverage | verbatim 每個字(繁簡正規化後)必須在此籤 source(raw ∪ page)字元集內 |
| A1b segment trace | 全部解曰的 ≥2 字片段必須是 per-slip source(raw_jie + 該籤自己的 page 解曰片段)的 substring,防拼裝/cross-slip |
| A2 chance isolation | chance 獨有字不得進 verbatim(chance 只能留 variants_or_notes 的 witness) |
| A3 uncertainty | UNRESOLVED:textual 必含「□」;structural(欄位格式/詩體)可無 □,note 需標 structural |
| A4 structure | 100 籤、每籤 2 筆、9 欄位齊全、layer_class 一致 |
| A5 encoding | 無 U+FFFD/mojibake/異常控制字符 |

輸入:`interpretation_layer.json` + `source_three_way.json`(raw_jie/page_ocr/chance_jie 三方)+ `per_slip_page_jie.json`(每籤自己的 page 解曰片段,由 `build_per_slip_page_jie.py` 從 page_ocr 提取、與 verbatim 匹配生成)。A1b 不直接使用 whole-page page_ocr(同頁另一籤會造成 cross-slip false pass),改用 per-slip page 片段;無可靠 page 片段(如 #21 七言詩)則只用 raw_jie。繁簡/異體經正規化(為↔爲、換↔换、虛↔虚、蟄↔蛰、晩↔晚 等)。A1 是 character-set coverage(只證字曾出現),A1b 才是詞序層的 substring traceability,兩者分工不宣稱 full traceability。

## UNRESOLVED 清單(解曰 20 籤)

| # | verbatim(缺字以□標記) |
|---|---|
| 1 | □□非□。□□時□。□音降□。報與君知。 |
| 4 | 五五念五。□龍跨虎。事雖勞心。於中有補。 |
| 5 | 望中心事。令可方求。百事營謀。正堪截□。 |
| 7 | 退身可得。進步難為。只宜守□。切莫高扳。 |
| 9 | 心中正直。原法寬□。天無私極。□空虛□。 |
| 10 | 機緣若遇。何事不成。春無限□。□似真□。 |
| 12 | □換得絲。是笑□哭。要見分明。是見為福。 |
| 13 | 因□得赦。病遇良醫。龍門得□。名顯□□。 |
| 14 | 從心無慮。遠達亨衢。道心自在。任□所如。 |
| 15 | 若得人怨。何事可伸。如言不信。到□勞心。 |
| 16 | 得處□□。損中有□。□□□凶。君子得吉。 |
| 17 | 心中不定。枉費看經。只是畫餅。□□□□。 |
| 20 | 佛神護佑。百事無虛。想平生事。到□勝初。 |
| 21 | 陰陽道合總由天。女嫁男婚豈偶然。但看龍蛇堪運動。熊羆葉夢喜團圓。 |
| 31 | 守己安靜。即是待他。時至必定。□全。 |
| 35 | 不須憂疑。自有□期。□前程。更換可宜。 |
| 44 | □求心事。如同□。要知勝負。先□□□。 |
| 48 | □□化□。諸禽不能。騰化時節。□□□□。 |
| 65 | 得止且止。知□割自。□肉痛本。一□。 |
| 77 | □夢說夢。聲名虛望。只好待時。□人□引。 |
| 96 | 這此福□。諸人皆現。可用誠心。福德即到。 |

## Sampled Direct Source-Image Verification(3 籤)

| # | 核對項目 | source image 結果 | 影響 |
|---|---|---|---|
| 1 | 解曰末句 | 「報與君知」(非 chance「先報君知」) | 以附錄二為準 |
| 4 | 解曰首句 | 「五五念五」(非 chance「淘沙成金」) | 確認附錄二原文 |
| 62 | 解曰全文 | 「諸事平穩,四方名顯,改換從新,凶存吉現」 | 三方一致,PROBABLE 確認 |

## Provenance / Reproducibility

- **source_locator**:`薛皓文2008附錄二 p{130–179}(籤 #N)`,每筆可回查原始掃描(appendix2_png 與逐籤 slip_pages/slip_XXX.png)
- **transcription_status**:解曰 80 PROBABLE(整頁 OCR + raw 兩獨立 OCR 一致或語義通順)+ 20 UNRESOLVED(三方分歧或字不通,留「□」不猜字);聖意 100 筆 verbatim 為附錄二 raw,直排順序亂,原始順序待逐籤 source image 還原
- **異體字**(如「𫝹」「冨」)原樣保存,不 canonicalize、不簡繁轉換(繁簡正規化僅用於 gate 判等,不改變 verbatim 文字)
- **chance substantive-variant**:逐條記錄 chance 與附錄二的實質差異於 `variants_or_notes`(chance 是艋舺版另一轉錄,非 production witness)

## 定位

本層是 archival / provenance interpretation layer。完整收錄不代表 UI 之後會全部展示;The Slip 仍是主體,Interpretation 是來源註腳。聖意原始直排順序與解曰 UNRESOLVED 20 籤,是下一輪 source image 逐籤核對的明確 backlog。
70 changes: 70 additions & 0 deletions data/corpora/guanyin/build_per_slip_page_jie.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# -*- coding: utf-8 -*-
"""從 page_ocr 提取每籤自己的 page 解曰片段(per-slip evidence)。

解決 cross-slip false pass:A1b 不得用 whole-page page_ocr 當 per-slip evidence。
本腳本從每籤 page_ocr 提取「該籤自己的解曰行」,去掉結尾標記,輸出 normalized 片段。
匹配原則:候選行與該籤 verbatim(去標點去□)共享字符最多;低分視為無可靠 page 片段(null)。

輸出 per_slip_page_jie.json:{slip_no: normalized片段 或 null}
"""
import json, sys, os

sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from validate_interpretation_layer import FAN2JIAN, normalize, strip_punct

HERE = os.path.dirname(os.path.abspath(__file__))
SRC = os.path.join(HERE, 'source_three_way.json')
LAYER = os.path.join(HERE, 'interpretation_layer.json')
OUT = os.path.join(HERE, 'per_slip_page_jie.json')

ENDS = ('此籤', '此鼔', '此籖', '此箴', '此跡', '此截', '此幾', '此果', '此乃', '此戰',
'此簽', '此藏', '此签', '此箝')

def candidates_from_page(poc):
out = []
for l in [x.strip() for x in poc.split('\n') if x.strip()]:
if not (8 <= len(l) <= 30):
continue
if any(k in l for k in ('龍山寺', '觀世音', '台北', '艋舺', '籤靈音世觀')):
continue
zx = l.find('之象')
if zx != -1 and zx < 8 and '凡事' in l: # 純象註行(「之象」在前 8 字)
continue
out.append(l)
return out

def strip_ends(s):
for e in ENDS:
if s.endswith(e):
return s[:-len(e)]
idx = s.rfind('之象')
if idx > len(s) // 2: # 「之象」在後半部 = 象註混入,截斷
return s[:idx]
return s

def main():
src = {x['slip_no']: x for x in json.load(open(SRC, encoding='utf-8'))}
layer = json.load(open(LAYER, encoding='utf-8'))
vt = {x['slip_no']: x['verbatim_text'] for x in layer['entries'] if x['field_type'] == '解曰'}

result = {}
for n in range(1, 101):
vt_n = normalize(strip_punct(vt[n]).replace('□', ''))
cand = candidates_from_page(src[n]['page_ocr'])
best, best_score = None, -1
for c in cand:
c_n = normalize(strip_punct(c))
score = len(set(c_n) & set(vt_n))
if score > best_score:
best_score, best = score, c_n
if best is not None and best_score >= 6:
result[str(n)] = strip_ends(best)
else:
result[str(n)] = None

json.dump(result, open(OUT, 'w', encoding='utf-8'), ensure_ascii=False, indent=2)
missing = [n for n, v in result.items() if v is None]
print(f'per_slip_page_jie.json 生成,缺失 {len(missing)} 籤: {sorted(map(int, missing))}')

if __name__ == '__main__':
main()
Loading
Loading