Skip to content

Commit abb2f15

Browse files
committed
feat(harness): let the decision model drive four judgement points
Compaction candidates, context modes, long-run steering, and long-term memory writes are now decided by the configured decision model when the matching strategy is set to ``decision``: - ``HARNESS_COMPACTION_STRATEGY=decision`` replaces the role based candidate rules with a content pre-filter plus a per-candidate "keep verbatim or summarize" judgement, because role labels cannot tell a tool result from the user's own text on real ADK traffic. - ``HARNESS_MODE_STRATEGY=decision`` decides the context mode blocks per invocation instead of matching keywords. - ``HARNESS_LONG_RUN_STRATEGY=decision`` only injects steering guidance while the run still looks unfinished, with the counter as a hard fallback after ``unconditional_after_model_calls``. - ``MEMORY_SAVE_STRATEGY=decision`` decides whether a turn holds something durable; a skipped turn keeps its cursor, so nothing is dropped. Every point defaults to its previous behaviour and degrades to it when the decision model is absent or fails. Change-Id: Id68bd9c71540e78a87638fc5c4735712cf7317a4
1 parent 5e75ce5 commit abb2f15

30 files changed

Lines changed: 2588 additions & 52 deletions

File tree

‎docs/content/docs/framework/memory/long-term/index.en.mdx‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -173,6 +173,16 @@ agent = Agent(
173173
```
174174

175175
To avoid frequent index re-initialization, VeADK exposes `MIN_MESSAGES_THRESHOLD` and `MIN_TIME_THRESHOLD` env vars to tune the save cadence: by default it saves after 10 accumulated events or a 60-second interval; additionally, when you switch `session_id` and start a new turn, VeADK saves the previous session to long-term memory.
176+
177+
The default is `MEMORY_SAVE_STRATEGY=threshold`, the two thresholds above. With `decision`, the configured decision model decides whether the turn holds something durable (a stated preference, fact, decision, or constraint), thresholded by `MEMORY_SAVE_WORTH_THRESHOLD` (default `0.5`):
178+
179+
| Judgement | Behaviour |
180+
| --- | --- |
181+
| Above the threshold | Saved immediately, without waiting for 10 events or 60 seconds |
182+
| Below the threshold | Skipped, and not saved merely because events accumulated |
183+
| Unavailable | Falls back to `MIN_MESSAGES_THRESHOLD` / `MIN_TIME_THRESHOLD` |
184+
185+
A skipped turn does not advance the save cursor, so the next accepted judgement writes those events together: nothing is lost. See the [environment variable reference](/references/configuration/environment-variables) for the decision model variables.
176186
## Auto-save Memory Policy
177187
Configure `auto_save_memory_policy` on `Agent` to decide which events are persisted by automatic long-term-memory saving. If omitted, it is equivalent to `"default"`.
178188
```python

‎docs/content/docs/framework/memory/long-term/index.mdx‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -173,6 +173,16 @@ agent = Agent(
173173
```
174174

175175
为避免索引被频繁初始化,VeADK 提供 `MIN_MESSAGES_THRESHOLD` 与 `MIN_TIME_THRESHOLD` 两个环境变量自定义保存周期:默认在累计 10 条 event 或间隔 60 秒时触发保存;此外,当切换 `session_id` 并发起新问答时,VeADK 会自动把上一个会话写入长期记忆。
176+
177+
默认保存策略是 `MEMORY_SAVE_STRATEGY=threshold`,即上面的双阈值规则。设置为 `decision` 后,改由已配置的判定模型判断"这一轮是否包含值得长期记住的内容"(陈述过的偏好、事实、决策、约束等),阈值由 `MEMORY_SAVE_WORTH_THRESHOLD` 控制,默认 `0.5`:
178+
179+
| 判定结果 | 行为 |
180+
| --- | --- |
181+
| 通过阈值 | 立即写入,不必等满 10 条 event 或 60 秒 |
182+
| 未通过阈值 | 跳过,不会仅因为攒够条数就写入 |
183+
| 判定不可用 | 回落到 `MIN_MESSAGES_THRESHOLD` / `MIN_TIME_THRESHOLD` |
184+
185+
被跳过的轮次不会推进保存游标,后续判定通过时会把这批 event 一起写入,因此不会丢数据。判定模型的环境变量见 [环境变量参考](/references/configuration/environment-variables)。
176186
## 自动保存记忆策略
177187
开发者在 `Agent` 上配置 `auto_save_memory_policy` 来控制自动保存长期记忆时哪些 event 会被写入;不配置时等价于 `"default"`。
178188
```python

‎docs/content/docs/references/configuration/environment-variables.en.mdx‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -76,6 +76,9 @@ Prefix `HARNESS_`, used to attach optional Harness plugins to HarnessApp Runtime
7676
| `HARNESS_MAX_TOOL_RESULT_CHARS` | Single tool-result compaction threshold, default `4000`. |
7777
| `HARNESS_VERIFIER_MODE` | Final-response verification mode, `observe` or `block`; default `observe`. |
7878
| `HARNESS_STORE_PATH` | Optional JSONL event store path; in-memory store is used when unset. |
79+
| `HARNESS_COMPACTION_STRATEGY` | Compaction candidate strategy, `builtin` or `decision`; default `builtin`. |
80+
| `HARNESS_LONG_RUN_STRATEGY` | Long-run steering strategy, `counter` or `decision`; default `counter`. |
81+
| `HARNESS_MODE_STRATEGY` | Context mode-block strategy, `keywords` or `decision`; default `keywords`. |
7982

8083
The `harness_enhance` block maps to these environment variables when deploying a HarnessApp Runtime. Prefer `harness.yaml` or `veadk agentkit invoke` flags for normal developer workflows; use environment variables for platform integration and container runtimes.
8184

‎docs/content/docs/references/configuration/environment-variables.mdx‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -76,6 +76,9 @@ volcengine:
7676
| `HARNESS_MAX_TOOL_RESULT_CHARS` | 单个工具结果压缩阈值,默认 `4000`。 |
7777
| `HARNESS_VERIFIER_MODE` | 最终回答校验模式,`observe` 或 `block`,默认 `observe`。 |
7878
| `HARNESS_STORE_PATH` | 可选 JSONL 事件存储路径;不设置时使用内存存储。 |
79+
| `HARNESS_COMPACTION_STRATEGY` | 工具结果压缩候选策略,`builtin` 或 `decision`,默认 `builtin`。 |
80+
| `HARNESS_LONG_RUN_STRATEGY` | 长任务收尾引导策略,`counter` 或 `decision`,默认 `counter`。 |
81+
| `HARNESS_MODE_STRATEGY` | 上下文模式块策略,`keywords` 或 `decision`,默认 `keywords`。 |
7982

8083
`harness_enhance` 配置块会在 HarnessApp Runtime 部署时映射为这些环境变量。推荐开发者优先通过 `harness.yaml` 或 `veadk agentkit invoke` 参数启用,环境变量适合平台集成和镜像运行时。
8184

‎docs/extensions/harness/README.md‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -173,6 +173,10 @@ export HARNESS_ENHANCE_ENABLED=true
173173
export HARNESS_ENHANCE_COMPONENTS=invocation_context,compactor,response_verification
174174
export HARNESS_COMPRESSION_PROVIDER=builtin
175175
export HARNESS_VERIFIER_MODE=observe
176+
# optional: let a decision model judge the built-in rules
177+
# export HARNESS_COMPACTION_STRATEGY=decision
178+
# export HARNESS_LONG_RUN_STRATEGY=decision
179+
# export HARNESS_MODE_STRATEGY=decision
176180
```
177181

178182
Equivalent YAML:
@@ -208,6 +212,21 @@ veadk agentkit invoke \
208212
| `HARNESS_MAX_TOOL_RESULT_CHARS` | `4000` | Tool-result compaction threshold. |
209213
| `HARNESS_VERIFIER_MODE` | `observe` | Verification behavior: `observe` or `block`. |
210214
| `HARNESS_STORE_PATH` | unset | Uses a JSONL event store when set. |
215+
| `HARNESS_COMPACTION_STRATEGY` | `builtin` | Compaction candidates: `builtin` or `decision`. |
216+
| `HARNESS_LONG_RUN_STRATEGY` | `counter` | Long-run steering: `counter` or `decision`. |
217+
| `HARNESS_MODE_STRATEGY` | `keywords` | Context mode blocks: `keywords` or `decision`. |
218+
219+
## Decision Model Strategies
220+
221+
The three `*_STRATEGY=decision` settings replace a rule with a judgement from
222+
the configured decision model. They need `DECISION_MODEL_ENABLED=true` and an
223+
API key; without one, each strategy keeps its rule and logs a warning.
224+
225+
| Strategy | Rule it replaces | Unavailable behaviour |
226+
| --- | --- | --- |
227+
| `HARNESS_COMPACTION_STRATEGY` | Role and size based compaction candidates | Builtin rules |
228+
| `HARNESS_LONG_RUN_STRATEGY` | Model-call counter | Counter, forced after the unconditional count |
229+
| `HARNESS_MODE_STRATEGY` | Precision and artifact keyword markers | Keyword markers |
211230

212231
## Compaction Providers
213232

‎docs/extensions/harness/README.zh.md‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -165,6 +165,10 @@ export HARNESS_ENHANCE_ENABLED=true
165165
export HARNESS_ENHANCE_COMPONENTS=invocation_context,compactor,response_verification
166166
export HARNESS_COMPRESSION_PROVIDER=builtin
167167
export HARNESS_VERIFIER_MODE=observe
168+
# 可选:把内置规则交给判定模型
169+
# export HARNESS_COMPACTION_STRATEGY=decision
170+
# export HARNESS_LONG_RUN_STRATEGY=decision
171+
# export HARNESS_MODE_STRATEGY=decision
168172
```
169173

170174
等价 YAML:
@@ -200,6 +204,19 @@ veadk agentkit invoke \
200204
| `HARNESS_MAX_TOOL_RESULT_CHARS` | `4000` | 工具结果压缩阈值。 |
201205
| `HARNESS_VERIFIER_MODE` | `observe` | 校验行为,支持 `observe` 或 `block`。 |
202206
| `HARNESS_STORE_PATH` | 未设置 | 设置后使用 JSONL event store。 |
207+
| `HARNESS_COMPACTION_STRATEGY` | `builtin` | 压缩候选策略:`builtin` 或 `decision`。 |
208+
| `HARNESS_LONG_RUN_STRATEGY` | `counter` | 长任务引导策略:`counter` 或 `decision`。 |
209+
| `HARNESS_MODE_STRATEGY` | `keywords` | 上下文模式块策略:`keywords` 或 `decision`。 |
210+
211+
## 判定模型策略
212+
213+
三个 `*_STRATEGY=decision` 开关把一条规则换成判定模型的判定结果,需要 `DECISION_MODEL_ENABLED=true` 与 API Key;没有配置时各自保留原规则并打印告警。
214+
215+
| 策略 | 被替代的规则 | 判定不可用时 |
216+
| --- | --- | --- |
217+
| `HARNESS_COMPACTION_STRATEGY` | 按角色和长度挑选压缩候选 | 内置规则 |
218+
| `HARNESS_LONG_RUN_STRATEGY` | 仅按模型调用次数计数 | 计数规则,超过强制次数后必定生效 |
219+
| `HARNESS_MODE_STRATEGY` | 精度/产物关键词匹配 | 关键词匹配 |
203220

204221
## 压缩 Provider
205222

Lines changed: 149 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,149 @@
1+
# Copyright (c) 2025 Beijing Volcano Engine Technology Co., Ltd. and/or its affiliates.
2+
#
3+
# Licensed under the Apache License, Version 2.0 (the "License");
4+
# you may not use this file except in compliance with the License.
5+
# You may obtain a copy of the License at
6+
#
7+
# http://www.apache.org/licenses/LICENSE-2.0
8+
#
9+
# Unless required by applicable law or agreed to in writing, software
10+
# distributed under the License is distributed on an "AS IS" BASIS,
11+
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12+
# See the License for the specific language governing permissions and
13+
# limitations under the License.
14+
15+
"""Harness judges against a real HTTP System One endpoint."""
16+
17+
from __future__ import annotations
18+
19+
import asyncio
20+
21+
import pytest
22+
23+
from veadk.extensions.decisions import (
24+
DecisionModelConfig,
25+
DecisionModelResponseError,
26+
DecisionExtension,
27+
)
28+
from veadk.extensions.harness.modules.tool_result_compactor import (
29+
DecisionCompactionJudge,
30+
ToolResultCompactor,
31+
ToolResultCompactorConfig,
32+
)
33+
from veadk.extensions.harness.schemas import CompressionRequest, ConversationMessage
34+
35+
from .fake_system_one import fake_system_one
36+
37+
38+
def _extension(base_url: str) -> DecisionExtension:
39+
return DecisionExtension(
40+
DecisionModelConfig(enabled=True, api_base=base_url, api_key="test-key")
41+
)
42+
43+
44+
def _scripted(
45+
probabilities: dict[int, float],
46+
) -> tuple[int, dict[str, str], dict[str, object]]:
47+
"""Build a response answering ``item_<index>`` with the given values."""
48+
answers = {
49+
f"item_{index}": {"type": "noul", "noul": value}
50+
for index, value in probabilities.items()
51+
}
52+
return 200, {}, {"model": "fake-system-one", "answers": answers, "usage": {}}
53+
54+
55+
def test_compaction_judge_batches_one_question_per_candidate() -> None:
56+
with fake_system_one() as server:
57+
judge = DecisionCompactionJudge(_extension(server.base_url))
58+
probabilities = asyncio.run(
59+
judge.aprotect(
60+
goal="rank the candidates by score",
61+
evidence={1: "x" * 5000, 3: "y" * 5000},
62+
)
63+
)
64+
65+
assert set(probabilities) == {1, 3}
66+
assert len(server.calls) == 1
67+
call = server.calls[0]
68+
assert call.model == "jev-latest"
69+
assert call.authorization == "Bearer test-key"
70+
assert sorted(call.questions) == ["item_1", "item_3"]
71+
assert call.questions["item_1"]["type"] == "noul"
72+
assert "rank the candidates by score" in call.state
73+
assert "item 1" in call.state and "item 3" in call.state
74+
75+
76+
def test_compaction_judge_parses_scripted_probabilities() -> None:
77+
with fake_system_one([_scripted({0: 0.95, 1: 0.05})]) as server:
78+
judge = DecisionCompactionJudge(_extension(server.base_url))
79+
probabilities = asyncio.run(
80+
judge.aprotect(goal="g", evidence={0: "a" * 4000, 1: "b" * 4000})
81+
)
82+
83+
assert probabilities[0] == pytest.approx(0.95)
84+
assert probabilities[1] == pytest.approx(0.05)
85+
86+
87+
def test_compaction_judge_rejects_a_partial_judgement() -> None:
88+
partial = (
89+
200,
90+
{},
91+
{"model": "fake", "answers": {"item_0": {"type": "noul", "noul": 0.9}}},
92+
)
93+
with fake_system_one([partial]) as server:
94+
judge = DecisionCompactionJudge(_extension(server.base_url))
95+
with pytest.raises(DecisionModelResponseError, match="no usable answer"):
96+
asyncio.run(
97+
judge.aprotect(goal="g", evidence={0: "a" * 4000, 1: "b" * 4000})
98+
)
99+
100+
101+
def test_compaction_judge_bounds_the_state_and_keeps_every_item() -> None:
102+
evidence = {index: "z" * 20000 for index in range(12)}
103+
with fake_system_one() as server:
104+
judge = DecisionCompactionJudge(
105+
_extension(server.base_url),
106+
max_state_chars=2400,
107+
max_evidence_chars=800,
108+
)
109+
asyncio.run(judge.aprotect(goal="g", evidence=evidence))
110+
111+
state = server.calls[0].state
112+
assert len(state) < 4000
113+
for index in evidence:
114+
assert f"item {index} (" in state
115+
116+
117+
def test_decision_strategy_end_to_end_keeps_evidence_and_fits() -> None:
118+
"""A real round trip through the client, the judge, and the policy."""
119+
120+
messages = [
121+
ConversationMessage(role="user", content="goal: rank the candidates"),
122+
ConversationMessage(role="user", content="tool_result: " + "x" * 9000),
123+
ConversationMessage(role="model", content="I inspected the table."),
124+
ConversationMessage(role="user", content="tool_result: " + "y" * 9000),
125+
ConversationMessage(role="model", content="One more check needed."),
126+
ConversationMessage(role="model", content="Now I will summarize."),
127+
ConversationMessage(role="user", content="go on"),
128+
]
129+
with fake_system_one([_scripted({1: 0.99, 3: 0.01})]) as server:
130+
compactor = ToolResultCompactor(
131+
ToolResultCompactorConfig(
132+
strategy="decision",
133+
max_context_chars=12000,
134+
summary_chars=400,
135+
),
136+
compaction_judge=DecisionCompactionJudge(_extension(server.base_url)),
137+
)
138+
result = asyncio.run(
139+
compactor.acompress_messages(
140+
CompressionRequest(messages=messages, max_context_chars=12000),
141+
goal="rank the candidates",
142+
)
143+
)
144+
145+
assert len(server.calls) == 1
146+
assert result.report.omitted_messages == 0
147+
assert result.report.compressed_chars <= 12000
148+
assert result.messages[1] == messages[1]
149+
assert result.messages[3] != messages[3]

0 commit comments

Comments
 (0)