Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -146,6 +146,14 @@ production D1からread-only取得した新しい5-node development / holdoutは

freeze後のdevelopmentでは、anchored 3 variantsがrelation MRRを`current`の0.5000から0.7500–1.0000へ改善しましたが、direct lookupとnegative-controlがともに退行しました。候補gate通過は0件だったためholdoutは開かず、既定strategyは`current_positive_additive`のままです。

## Anchored fusion calibration experiment

Issue #15で分離したentry anchorとedge-only graph signalは維持したまま、graph尺度とfinal fusionだけを比較します。graph normalizationは`max`、rawの`none`、`l1_mass`を選択でき、final fusionはlinearとpositive graph nodeだけを順位付けするbottom-centered weighted RRFを選択できます。

production D1からread-only取得した新しい3-node development / holdoutは、既存7 fixturesの50 unique doc pathsおよび相互間から分離しています。固定6 variants、fusion formula、個別case non-regression gate、one-time holdout停止規則は[Anchored fusion calibration experiment](docs/anchored-fusion-calibration-experiment.md)を参照してください。

freeze後のdevelopmentでは、unscaled linearとbalanced RRFがrelation MRRを0.4167から0.6667へ改善しましたがdirect / negative-controlを1.0000から0.7500へ退行させました。conservative linear、L1 mass、conservative RRFはcontrolsを維持した一方relationを改善しませんでした。候補0件のためholdoutは未開封で、既定strategyは`current_positive_additive`のままです。

## Public API

```python
Expand Down Expand Up @@ -212,6 +220,8 @@ MCP 対応 AI との接続は、コアへ MCP SDK を追加せず、同一 repos
- final score
- seed node
- zero-hop / graph path種別
- entry / positive graph rank
- entry / graph fusion componentとfusion strategy
- path contribution
- path 上の edge type、weight、factuality

Expand Down
108 changes: 108 additions & 0 deletions docs/anchored-fusion-calibration-experiment.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
# Anchored fusion calibration experiment

## 目的

前回のanchored local experimentはentry anchor invariant、edge-only graph signal、relation path、feedback isolationを満たし、relation MRRを改善した一方でdirect lookupとnegative-controlを退行させた。本実験はactivation dynamicsとBM25/dense entryを変更せず、graph normalizationとfinal fusionだけを切り分ける。

既定は`current_positive_additive`のままとし、未観測holdoutで固定gateを通過した場合だけ変更候補とする。

## 凍結artifact

- manifest: `tests/fixtures/d1_liplus_fusion_calibration_experiment.manifest.json`
- development fixture / gold / provenance: `d1_liplus_fusion_calibration_development.*.json`
- holdout fixture / gold / provenance: `d1_liplus_fusion_calibration_holdout.*.json`
- contamination audit: `d1_liplus_fusion_calibration.contamination.json`

developmentはsubagent parallel width cluster、holdoutはSheepdog Engineering clusterで、各3 nodes / 2 edgesのweakly-connected componentである。production D1へSELECT / WITHだけを実行し、全queryで`rows_written=0`、`changes=0`、`changed_db=false`を検証する。

既存7 fixturesのunionである50 unique doc pathsをdenylistとする。新旧および新split相互のdoc path、node ID、source URL、normalized query、expected node、relation endpoint重複を拒否する。prior goldとprior resultはauditへ読み込まない。

3-node componentは全既存pathを除外したproduction D1で確保できる最大の相互disjoint connected componentである。rank指標が粗くなるため、cohort MRRだけでなく個別case rankもgateに含める。

## normalizationとfusion

Graph normalization:

- `max`: positive graph activationを最大値で割る。最大値0なら全0
- `none`: raw positive graph activation
- `l1_mass`:各activationをpositive activation総和で割る。総和0なら全0

Linear fusionは正規化したweightでentry anchorとgraphを加算する。weight合計が1の本実験では`entry_weight * entry + graph_weight * graph`と同一である。

Weighted RRFはentry全nodeをscore降順・node ID昇順で順位付けし、graphはraw activationが正のnodeだけを同じ規則で順位付けする。graph activationが0のnodeはgraph rankを持たず、graph componentを0とする。

graph-positive node数を`P`として、bottom-centered式を使う。

```text
entry_component = entry_weight / (k + entry_rank)
graph_component = graph_weight * (1 / (k + graph_rank) - 1 / (k + P + 1))
final = entry_component + graph_component
```

`k=60`。bottom subtractionはpositive graph集合の仮想最下位rankを0基準にし、graph channelが欠けるentry seedへ過大なpenaltyを与えないために固定する。

各hitはraw / normalized score、entry / graph rank、両fusion component、final、normalization、fusion strategyを保持する。final降順・node ID昇順をresultから再計算できなければgate不合格とする。

## 固定variants

最大数と実数を6に固定し、parameter gridを追加しない。

1. `current`: current positive-additive、max、linear、0.55 / 0.45
2. `anchored-local-unscaled`: anchored local、none、linear、0.55 / 0.45
3. `anchored-linear-conservative`: anchored local、max、linear、0.80 / 0.20
4. `anchored-linear-mass`: anchored local、l1_mass、linear、0.70 / 0.30
5. `anchored-rrf-conservative`: anchored local、bottom-centered RRF、0.80 / 0.20、k=60
6. `anchored-rrf-balanced`: anchored local、bottom-centered RRF、0.65 / 0.35、k=60

## development選択規則

candidateは次をすべて満たす必要がある。

- relation MRRが`current`より厳密に高い
- relation caseを最低1件個別rank改善する
- direct lookup / negative-control MRRが`current`から退行しない
- 全direct / negative-control caseの個別rankが`current`から退行しない
- 全relation pathが固定endpoint / edge typeと一致する
- success feedbackが最低1 edgeをcreditし、uncredited edgeと非対象rankを変更しない
- entry anchorがcompetition前後で不変である
- graph signalとgraph rankがzero-hop pathを含まない
- traceから全final componentとorderingを再計算できる

複数候補はrelation MRR、worst-cohort MRR、個別改善case数、平均expansions、structural complexity、variant IDの順で一意に選ぶ。候補がなければholdoutを開かない。

## holdout停止規則

development候補がある場合だけ、未観測holdoutで`current`と選択候補を一度評価する。採用には同じgateを要求する。一つでも失敗すればdefaultを維持する。

fixture、gold、provenance、audit、manifest、normalization、weights、RRF k、threshold、selection rule、stop ruleはfreeze commitをpushするまでrunnerへ渡さない。development / holdout resultは既存fileを上書きせず、観測後の変更と再実行を禁止する。

## 実行手順

freeze commitのpush後にdevelopmentを一度だけ実行する。

```powershell
uv run python tools/run_dynamics_experiment.py development `
--manifest tests/fixtures/d1_liplus_fusion_calibration_experiment.manifest.json `
--output tests/fixtures/d1_liplus_fusion_calibration_experiment.development.result.json
```

候補がある場合だけholdoutを一度実行する。候補がなければholdout resultを作成しない。

## 観測結果

fixture / gold / provenance / audit / manifest / implementation / tests / 規則をfreeze commit `f801264`としてpushした後、developmentを一度だけ実行した。

| variant | direct MRR | relation MRR | negative MRR | gate |
| --- | ---: | ---: | ---: | --- |
| `current` | 1.0000 | 0.4167 | 1.0000 | baseline |
| `anchored-local-unscaled` | 0.7500 | 0.6667 | 0.7500 | control cohort / 個別rank退行 |
| `anchored-linear-conservative` | 1.0000 | 0.4167 | 1.0000 | relation非改善 |
| `anchored-linear-mass` | 1.0000 | 0.4167 | 1.0000 | relation非改善 |
| `anchored-rrf-conservative` | 1.0000 | 0.4167 | 1.0000 | relation非改善 |
| `anchored-rrf-balanced` | 0.7500 | 0.6667 | 0.7500 | control cohort / 個別rank退行 |

全anchored variantsでrelation path、feedback isolation、entry anchor invariant、edge-only graph signal、formula / ordering再計算が成立した。

unscaled linearとbalanced RRFはrelationを個別改善したが、direct / negative-controlのcohort MRRと個別rankを退行させた。conservative linear、L1 mass、conservative RRFは全controlを維持したがrelationを個別改善しなかった。固定gateを通る候補は0件だった。

selectionは`current`、理由は`no_fusion_variant_passed_frozen_gate`である。停止規則に従ってholdoutを開封せず、holdout resultを作成しない。defaultは`current_positive_additive`のままとする。
7 changes: 7 additions & 0 deletions docs/requirements.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,12 @@
47. anchored hybrid experimentは`current`、`bm25-only`、BM25+現行graph、BM25+dense anchorのlocal/query local、BM25 anchorのlocalを合わせた6 variantsだけを比較する。
48. development候補はrelation MRRがcurrentを厳密に上回り、direct / negative-controlが退行せず、relation path、feedback isolation、anchor invariant、edge-only graph signalをすべて満たす場合だけ選択する。
49. 候補がある場合だけ未観測holdoutで`current`、`bm25-only`、選択候補を一度評価し、同じgateを通過した場合だけdefault変更候補とする。
50. graph activationは`max`、`none`、`l1_mass`を一般設定として選択でき、zero totalを決定論的に全0へ変換できる。
51. final fusionは既存linearに加え、entry rankとpositive graph nodeだけのgraph rankを使うbottom-centered weighted RRFを選択できる。
52. 各traceはentry / graph rank、entry / graph fusion component、normalization、fusion strategy、RRF k、positive graph node数を保持し、final orderingを機械的に再計算できる。
53. fusion calibration experimentはproduction D1の新しい3-node development / holdoutを、既存7 fixturesの50 unique doc pathsおよび相互間から分離し、provenance、balanced gold、contamination audit、6 variantsを結果観測前に固定する。
54. development候補はrelation MRRをcurrentから厳密に改善し、少なくとも1 relation caseを個別改善し、direct / negative-controlのcohort MRRと全個別rankを退行させず、path、feedback、anchor、edge-only graph、formula auditを満たす場合だけ選択する。
55. development候補がある場合だけ未観測holdoutでcurrentと選択候補を一度評価し、同じgateを通過した場合だけdefault変更候補とする。

## 4. Constraints

Expand All @@ -88,3 +94,4 @@
- [Neural dynamics experiment](neural-dynamics-experiment.md) が development / holdout 分離、固定探索空間、候補選択、単一 holdout 開封、停止規則を定義する。
- [Local recurrent competition experiment](neural-dynamics-local-competition-experiment.md) が新規D1 subgraph、contamination audit、query / path ablation、二baseline gateを定義する。
- [Anchored BM25 and graph hybrid experiment](anchored-bm25-graph-hybrid-experiment.md) がzero-hop anchor、edge-only graph signal、BM25 ablation、新規D1 split、単一holdout gateを定義する。
- [Anchored fusion calibration experiment](anchored-fusion-calibration-experiment.md) がgraph normalization、bottom-centered RRF、新規D1 split、個別case gateを定義する。
124 changes: 111 additions & 13 deletions src/neuron_graph_rag/engine.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,9 @@ class EngineConfig:
max_active_paths_per_node: int = 4
use_dense_retrieval: bool = True
use_graph_propagation: bool = True
graph_normalization: str = "max"
final_fusion_strategy: str = "linear"
rrf_k: int = 60

def __post_init__(self) -> None:
for name in (
Expand Down Expand Up @@ -106,6 +109,12 @@ def __post_init__(self) -> None:
raise ValueError("dense_weight must be zero when dense retrieval is disabled")
if not self.use_graph_propagation and self.graph_weight != 0.0:
raise ValueError("graph_weight must be zero when graph propagation is disabled")
if self.graph_normalization not in {"max", "none", "l1_mass"}:
raise ValueError("Unknown graph_normalization")
if self.final_fusion_strategy not in {"linear", "rrf"}:
raise ValueError("Unknown final_fusion_strategy")
if self.rrf_k < 1:
raise ValueError("rrf_k must be positive")


class NeuronGraphRAG:
Expand Down Expand Up @@ -244,24 +253,36 @@ def search(
"active_path_count": 0,
"competition_sets": [],
}
normalized_activation = normalize_scores(
{node.node_id: graph_activation.get(node.node_id, 0.0) for node in nodes}
graph_values = {
node.node_id: graph_activation.get(node.node_id, 0.0) for node in nodes
}
normalized_activation = self._normalize_graph_activation(
graph_values, self.config.graph_normalization
)
self._record_activation(graph_activation, timestamp)

hits = [
SearchHit(
entry_ranks = self._rank_positive_or_all(entry, positive_only=False)
graph_ranks = self._rank_positive_or_all(
graph_values, positive_only=True
)
positive_graph_count = len(graph_ranks)

hits = []
for node in nodes:
entry_component, graph_component = self._fusion_components(
entry=entry[node.node_id],
normalized_graph=normalized_activation[node.node_id],
entry_rank=entry_ranks[node.node_id],
graph_rank=graph_ranks.get(node.node_id),
positive_graph_count=positive_graph_count,
)
hits.append(SearchHit(
node=node,
sparse_score=sparse[node.node_id],
dense_score=dense[node.node_id],
entry_score=entry[node.node_id],
graph_activation=graph_activation.get(node.node_id, 0.0),
final_score=self._weighted_average(
entry[node.node_id],
normalized_activation[node.node_id],
self.config.entry_weight,
self.config.graph_weight,
),
final_score=entry_component + graph_component,
paths=tuple(
sorted(
paths.get(node.node_id, ()),
Expand All @@ -273,9 +294,13 @@ def search(
normalized_graph_activation=normalized_activation[node.node_id],
entry_anchor_before_competition=entry[node.node_id],
entry_anchor_after_competition=entry[node.node_id],
)
for node in nodes
]
entry_rank=entry_ranks[node.node_id],
graph_rank=graph_ranks.get(node.node_id),
entry_fusion_component=entry_component,
graph_fusion_component=graph_component,
final_fusion_strategy=self.config.final_fusion_strategy,
graph_normalization=self.config.graph_normalization,
))
hits.sort(key=lambda hit: (-hit.final_score, hit.node.node_id))
selected_hits = tuple(hits[:limit])
trace_id = uuid.uuid4().hex
Expand Down Expand Up @@ -311,6 +336,15 @@ def search(
for node_paths in paths.values()
for path in node_paths
),
"graph_normalization": self.config.graph_normalization,
"final_fusion_strategy": self.config.final_fusion_strategy,
"rrf_k": self.config.rrf_k,
"positive_graph_node_count": positive_graph_count,
"final_order_recomputable": all(
hit.final_score
== hit.entry_fusion_component + hit.graph_fusion_component
for hit in hits
),
}
)
return SearchTrace(
Expand Down Expand Up @@ -439,6 +473,70 @@ def _weighted_average(
total_weight = left_weight + right_weight
return (left * left_weight + right * right_weight) / total_weight

@staticmethod
def _normalize_graph_activation(
scores: dict[str, float], strategy: str
) -> dict[str, float]:
positive = {key: max(0.0, value) for key, value in scores.items()}
if strategy == "none":
return positive
if strategy == "max":
return normalize_scores(positive)
if strategy == "l1_mass":
total = sum(positive.values())
if total <= 0.0:
return {key: 0.0 for key in positive}
return {key: value / total for key, value in positive.items()}
raise ValueError(f"Unknown graph normalization: {strategy}")

@staticmethod
def _rank_positive_or_all(
scores: dict[str, float], *, positive_only: bool
) -> dict[str, int]:
ordered = sorted(
(
(node_id, score)
for node_id, score in scores.items()
if not positive_only or score > 0.0
),
key=lambda item: (-item[1], item[0]),
)
return {
node_id: rank
for rank, (node_id, _) in enumerate(ordered, start=1)
}

def _fusion_components(
self,
*,
entry: float,
normalized_graph: float,
entry_rank: int,
graph_rank: int | None,
positive_graph_count: int,
) -> tuple[float, float]:
if self.config.final_fusion_strategy == "linear":
total_weight = self.config.entry_weight + self.config.graph_weight
return (
entry * self.config.entry_weight / total_weight,
normalized_graph * self.config.graph_weight / total_weight,
)
entry_component = self.config.entry_weight / (
self.config.rrf_k + entry_rank
)
graph_component = 0.0
if graph_rank is not None:
graph_component = self.config.graph_weight * (
1.0 / (self.config.rrf_k + graph_rank)
- 1.0
/ (
self.config.rrf_k
+ positive_graph_count
+ 1
)
)
return entry_component, graph_component

@staticmethod
def _timestamp(value: datetime | float | None) -> float:
if value is None:
Expand Down
Loading