Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,6 +101,8 @@ GitHub webhooks + GitHub API

成功検索は既定で UTC trace を保存し、資料の版を指す `source_id` と現在 activation を返します。`memory_history` で履歴・利用段階・取消し audit を読み、`record_source_use` で selected→validated→used、`record_outcome` で confirmed または receipt を指定した corrected/rolled_back を記録します。検索だけでは used/confirmed になりません。保存不可の結果には trace がなく feedback 不可です。半減期3600秒、activation は各 channel 上限10(total20)、graph strength 上限5です。成功検索に DO request と保存コストが加わります。

識別できない資料が混ざっても正常資料の trace は保存します。`memory_recording` が部分記録の件数・理由を返し、履歴の `settings.memory_recording` に残ります。未解決 row は `feedback_available:false` で source_id を持ちません。全資料未解決は trace を発行せず、正常な0件検索は空 trace を保存します。検証不可の graph path は強化されず、正常な終点資料の利用記録は可能です。検索結果・順序は維持します。

live inline doc/wiki 本文は返した本文の SHA-256 で版を識別し、`github_live` provenance を持ちます。stored fetch と graph 本文は index snapshot の版で、索引 timestamp が同じでも live 本文の変更は別 source_id になります。

[仕様・移行・合成実演](docs/2-feedback-memory.ja.md) を参照してください。0008 を Worker deploy より先に適用し、新 schema JSON を bridge artifact に同梱します。検索品質向上は未評価です。
Expand Down
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,6 +101,8 @@ This MCP server exposes `search` and three private feedback tools. All retrieval

Successful retrieval writes a UTC trace by default, returning version-specific `source_id` and current activation. `memory_history` reads history, stages and reversal audit; `record_source_use` records selected→validated→used; `record_outcome` confirms a saved path or corrects/rolls back one named receipt. Retrieval alone is never usage or confirmation. Memory failure emits no usable trace. Half-life is 3600 seconds, each activation channel caps at 10 (total 20), and relation strength caps at 5. Successful retrieval adds a DO request and storage cost.

Unresolved identities exclude only affected sources; valid sources still save a trace. `memory_recording` reports partial counts/reasons and persists in history settings. Unresolved rows have feedback_available:false and no source_id. All-unresolved results save no trace; genuine zero-hit retrieval saves an empty trace. Unverifiable graph paths cannot receive credit, while resolved terminals can record usage. Retrieval rows and order remain intact.

Live inline doc/wiki bodies use the returned-text SHA-256 version and `github_live` provenance. Stored fetch and graph bodies refer to the index snapshot; changed live text has another source ID even with the same index timestamp.

See the [specification, migration and synthetic lifecycle](docs/2-feedback-memory.md). Apply 0008 before Worker deployment and include the new schema JSON in bridge artifacts. Retrieval quality improvement has not been evaluated.
Expand Down
4 changes: 4 additions & 0 deletions docs/0-requirements.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -666,6 +666,10 @@ embedding が失敗した record は incomplete と分かる形で残し、次

新 API は `memory_history`、`record_source_use`、`record_outcome`。設定値、上限、ID/version、batch/idempotency、ledger、error、0008先行移行、既存行 backfill と artifact の契約は [2-feedback-memory.ja.md](2-feedback-memory.ja.md) に定義する。会話/本文/credential の複製は行わず、品質向上の主張は別評価を必要とする。

Issue #263: canonical identity の失敗は資料単位で除外し、正常資料は既存の原子的保存に渡す。検索結果と順序は維持し、部分記録の状態・除外件数・理由を応答と保存 settings に残す。全資料未解決は正常な0件検索と区別する。未解決資料に source_id を発行せず、feedback を不可とする。経路の起点・中間・終点を確認できない graph path は強化対象から除外する。保存失敗と実際の部分取得は引き続き trace を発行しない。

Issue #263 の調査で再現した issue_comment.deleted の部分削除: Vectorize とFTSの両方が削除成功するまでDO canonicalを保持し、片側失敗とDO削除失敗は503を返す。自動再試行は追加しない。対象の残存1件との因果は未確認であり、[調査記録](3-unresolved-source-investigation.ja.md) に事実と仮説を分けて保存する。本番DB変更・一括補修・版番号更新・releaseはこの修正に含めない。

#259 の版契約: live inline doc/wiki は返した本文の SHA-256 を version とし、index snapshot と provenance を分ける。同じ索引 timestamp でも live 本文更新は別 source_id にする。本文を memory に複製しない。

## リリース metadata の一致(Issue #261)
Expand Down
4 changes: 4 additions & 0 deletions docs/0-requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -671,6 +671,10 @@ Implement retrieval history, selected/validated/used records, retrieved/usage ac

The APIs are `memory_history`, `record_source_use`, and `record_outcome`. [2-feedback-memory.md](2-feedback-memory.md) owns constants, bounds, source version identity, batch/idempotency, ledger, errors, migration 0008-before-deployment, historical backfill and bridge artifacts. Do not duplicate conversations, bodies or credentials. Claims of retrieval quality improvement require separate evaluation.

Issue #263: exclude canonical identity failures per source and pass valid sources to the existing atomic save. Preserve retrieval rows and order; return and persist partial-recording status, exclusion counts and reasons in settings. Distinguish all-unresolved from a genuine zero-hit search. Unresolved rows receive no source_id and cannot accept feedback. Exclude graph paths from reinforcement unless the origin, intermediates and terminal can be verified. Storage failures and incomplete retrieval continue to emit no trace.

Issue #263 investigation reproduces partial issue_comment.deleted teardown: retain DO canonical identity until both Vectorize and FTS deletions succeed. Index or DO deletion failure returns 503; no automatic retry is added. Causality for the remaining production row is unverified; [investigation](3-unresolved-source-investigation.ja.md) separates evidence from hypotheses. Production DB mutation, bulk repair, version updates and releases are outside this fix.

Issue #259 version contract: live inline doc/wiki uses the returned-text SHA-256 version and provenance separate from the index snapshot. Changed live text gets a different source ID even at the same index timestamp. Do not duplicate the body in memory.

## Release metadata consistency (Issue #261)
Expand Down
14 changes: 12 additions & 2 deletions docs/2-feedback-memory.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,17 @@ principal は検証済み MCP OAuth props の numeric GitHub user ID からサ

limit は1..50、trace の資料 snapshot は300件、保存入力は300,000文字、query は4096文字、metadata string は256文字、usage batch は100 entry、reason/key は1000/128文字まで。SQL は owner/cursor index と有界 page を使う。全履歴をまとめて読む API はない。

未成立・失敗を成功 trace として保存しない。一部 scan source、sparse retrieval、graph expansion の失敗、canonical identity 未解決、memory の原子的書込失敗では検索結果を返せるが、`memory_unavailable:true`、`feedback_available:false` とし、**trace_id を発行しない**。この結果への feedback は不可。拒否 error は段階違反、key conflict、confirmation 等の安全な code を返し、query/handle を error や log へ転記しない。検索の error log は固定文言にする。
未成立・失敗を成功 trace として保存しない。一部 scan source、sparse retrieval、graph expansion の失敗と memory の原子的書込失敗では検索結果を返せるが、`memory_unavailable:true`、`feedback_available:false` とし、**trace_id を発行しない**。`memory_error` はそれぞれ `retrieval_incomplete` / `memory_save_failed`。拒否 error は段階違反、key conflict、confirmation 等の安全な code を返し、query/handle を error や log へ転記しない。検索の error log は固定文言にする。

### 資料単位の部分記録(Issue #263)

canonical identity または live version が未解決の資料は、検索結果・順序・本文を維持して、その資料だけ保存から除外する。正常資料があれば既存 transaction で原子的に保存し `trace_id` と全体の `feedback_available:true` を返す。未解決 row は `feedback_available:false`、`memory_exclusion_reason:canonical_identity_unavailable` または `live_version_unavailable` を持ち、source_id/provenance/activation を付けない。`results`、`graph_results`、両軸の `same_entity.others` に適用する。

応答の `memory_recording` は `status:complete|partial|unresolved`、`recorded_sources`(保存した distinct source 数)、`excluded_sources`(除外した返却 row 数)、`exclusions:[{location,reason}]` を返す。同じ未解決資料が複数箇所にあれば各 row を数える。location は `results[0].same_entity.others[1]` 等の返却位置のみ。本文・タイトル・vector ID を除外記録へコピーしない。同じ object を `settings.memory_recording` として保存し、history 一覧・詳細の両方で再読できる。古い trace の settings にこの field が無い場合は従来の記録。

全資料未解決は `status:unresolved`、`memory_error:all_sources_unresolved`、memory_unavailable/feedback不可、trace無し。正常な0件検索は `status:complete`、件数0の空 trace を保存する。識別以外の例外とDB書込失敗は資料単位の除外として扱わない。

graph path の起点・中間・終点は実際に探索した vector ID 列から索引 metadata を読む。返却枠外・dangling・identity 未解決の中間も検証し、repo と directional mention の両端 slug を照合する。確認不可の path は保存 snapshot で `path:[]` とし、返却 graph_path は維持する。終点自体が正常なら利用段階は記録可能だが `graph_feedback_available:false` と `graph_feedback_exclusion_reason:unverifiable_graph_path` を返し、confirmed は `no_graph_path` で非適用。`excluded_graph_paths` / `graph_exclusions` を memory_recording に保存する。確認できた他の path の強化・receipt・取消しは継続する。内部 node snapshot は返却・settings・history に含めない。graph metadata 読取は返却 path 当たり最大3 node の集合に限定し、保存形式の移行は不要。

## 利用段階と retry

Expand Down Expand Up @@ -72,7 +82,7 @@ npx wrangler d1 execute github-rag-fts --remote --command "SELECT type,COUNT(*)

deploy 後の認可された `POST /admin/backfill-source-identities?repo=owner/repo&limit=50[&cursor=TYPE:ID]` は、既存 DO の canonical event ID から過去の unchanged 行を補修する。既存管理 credential は非公開 header で渡し、URL / artifact / shell history に token を載せない。next_cursor を渡して done:true まで進める。limit は1..100、1 page は最大3つの有界 DO read と一つの原子的 D1 batch。再開・再実行は安全で、embedding / GitHub 本文再取得 / index reset は不要。

通常の unchanged ingest も hash skip の前に ID を補修する。DO に canonical event が無い行は未解決のままなので、別の再取り込み前に欠損を調べる。未解決 source を含む検索は memory を fail-closed にする。backfill は private trace/credit を変更しない。
通常の unchanged ingest も hash skip の前に ID を補修する。DO に canonical event が無い行は未解決のままなので、別の再取り込み前に欠損を調べる。未解決 source 自体は feedback 不可とし、正常 source は上記の部分記録を適用する。backfill は private trace/credit を変更しない。旧ID欠落と残存1件の調査、再現した削除経路の修正は [調査記録](3-unresolved-source-investigation.ja.md) を参照。

bridge には `server/search-schema.json` と `server/memory-tools.json` を同梱する。`node scripts/generate-tool-contracts.mjs` で Worker Zod contract から生成し、`node scripts/check-schema-drift.mjs` が全 tool の実 protocol と nested 入出力 schema / bounds / defaults / annotations の完全一致を確認する。npm / mcpb の両 artifact にこの JSON が必要。stdio client の tool discovery には更新 bridge の公開が必要。公開 version と publication は後続の運用段階で決める。

Expand Down
12 changes: 10 additions & 2 deletions docs/2-feedback-memory.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,15 @@ The server derives `github:<numeric GitHub user ID>` from verified MCP OAuth pro

`memory_history {limit:10}` returns newest trace summaries and `next_cursor`; pass that cursor for the next page. `memory_history {trace_id,limit:10}` returns sources, current stages/activation, settings and oldest-first receipt audit. Its `next_cursor` pages **receipts within that trace**, not trace summaries. Cursors belong to the authenticated principal and, for details, the trace. Limits are 1..50; source snapshots cap at 300 per trace, saved trace input at 300,000 characters, queries at 4096 characters, metadata strings at 256, usage batches at 100 entries, and reason/idempotency handles at 1000/128 characters. SQL reads use owner/cursor indexes and bounded pages; no full-history fetch is used.

A failed retrieval is not recorded as successful. If retrieval is partial (a scan surface, sparse retrieval, or graph expansion failed), identity cannot be resolved, or the atomic memory write fails, retrieval can still return with `memory_unavailable: true`, `feedback_available: false` and **no `trace_id`**. It cannot accept feedback. Memory rejection codes describe ownership, stage, key conflict and confirmation errors without echoing the submitted query or handles. Retrieval error logging emits fixed messages.
A failed retrieval is not recorded as successful. If retrieval is partial (a scan surface, sparse retrieval, or graph expansion failed), or the atomic memory write fails, retrieval can still return with `memory_unavailable: true`, `feedback_available: false` and **no `trace_id`**. `memory_error` distinguishes `retrieval_incomplete` and `memory_save_failed`. Memory rejection codes describe ownership, stage, key conflict and confirmation errors without echoing the submitted query or handles. Retrieval error logging emits fixed messages.

### Partial source recording (Issue #263)

Canonical identity and live-version failures exclude only the affected rows from storage; retrieval rows, order and text remain intact. Valid sources commit atomically with a trace and overall `feedback_available:true`. Unresolved rows return `feedback_available:false` and `memory_exclusion_reason:canonical_identity_unavailable|live_version_unavailable`, without source_id/provenance/activation. This applies to results, graph_results and their same_entity.others members.

`memory_recording` returns status complete/partial/unresolved, recorded_sources (distinct saved sources), excluded_sources (excluded returned occurrences), and exclusions [{location,reason}]. Locations name response positions only, e.g. results[0].same_entity.others[1]; no text, title or vector ID is copied. The same object persists as settings.memory_recording in history summaries and details. Old traces lack this field. All-unresolved results emit memory_error:all_sources_unresolved, memory_unavailable and no trace. Genuine zero-hit retrieval saves an empty complete trace. Unexpected exceptions and storage failures are not classified as row exclusions.

Graph origins, intermediates and terminals are checked against indexed metadata using the actual traversed vector IDs, including nodes omitted from the output and dangling nodes. Verify canonical identity, repository and directional mention endpoint slugs. An unverifiable path saves path:[] while retaining the returned graph_path. A resolved terminal can record usage but returns graph_feedback_available:false and graph_feedback_exclusion_reason:unverifiable_graph_path; confirmation yields no_graph_path without credit. memory_recording retains excluded_graph_paths and graph_exclusions. Other verified paths keep confirmation receipts and reversal. Internal node snapshots are removed before response/storage; graph metadata reads cover the union of at most three nodes per returned path. No storage migration is needed.

## Usage and idempotency

Expand Down Expand Up @@ -62,7 +70,7 @@ npx wrangler d1 execute github-rag-fts --remote --command "PRAGMA table_info(sea
npx wrangler d1 execute github-rag-fts --remote --command "SELECT type,COUNT(*) AS unresolved FROM search_docs WHERE (type IN ('issue_comment','pr_review_comment') AND comment_id=0) OR (type='pr_review' AND review_id=0) GROUP BY type"
```

After deployment, authorized `POST /admin/backfill-source-identities?repo=owner/repo&limit=50[&cursor=TYPE:ID]` repairs historical unchanged rows from existing canonical DO events. Supply the existing administrative credential through a private header; no token belongs in a URL, artifact or shell history. Follow `next_cursor` until `done:true`. Limit 1..100, at most three bounded DO reads and one atomic D1 batch per page. Restarting is safe; no embeddings, GitHub body refetch or index reset is needed. Normal unchanged comment/review ingest also repairs IDs before its hash-skip return. Rows whose canonical event is absent from the DO remain unresolved; inspect that gap before any separate reingest. A trace containing unresolved source identity fails memory closed. Backfill does not rewrite private traces/credits.
After deployment, authorized `POST /admin/backfill-source-identities?repo=owner/repo&limit=50[&cursor=TYPE:ID]` repairs historical unchanged rows from existing canonical DO events. Supply the existing administrative credential through a private header; no token belongs in a URL, artifact or shell history. Follow `next_cursor` until `done:true`. Limit 1..100, at most three bounded DO reads and one atomic D1 batch per page. Restarting is safe; no embeddings, GitHub body refetch or index reset is needed. Normal unchanged comment/review ingest also repairs IDs before its hash-skip return. Rows whose canonical event is absent from the DO remain unresolved; inspect that gap before any separate reingest. Unresolved sources cannot accept feedback; valid sources use partial recording above. Backfill does not rewrite private traces/credits. See the [investigation](3-unresolved-source-investigation.ja.md) for historical ID omission, the remaining row, and the reproduced deletion-path fix.

The bridge includes `server/search-schema.json` and `server/memory-tools.json`. `node scripts/generate-tool-contracts.mjs` regenerates them from Worker Zod schemas; `node scripts/check-schema-drift.mjs` compares all shipped tools' exact nested input/output contracts, bounds, defaults and annotations against the actual protocol. Package both JSON files in npm/mcpb artifacts. Worker deployment enables direct-client tools; the published bridge must include the new artifacts for stdio clients to discover them. Public version selection and publication are subsequent operations.

Expand Down
Loading
Loading