Skip to content

Commit 4b2639b

Browse files
committed
feat(harness): judge the advertised skills and the transfer target
Two more places decide with nothing but wording. The skills callback advertises every loaded skill in the agent instruction, so a large library spends prompt budget on skills the request will never use, and the model picks from a list nobody narrowed. Routing a request to a sub-agent has the same shape: the model reads the agent descriptions and calls ``transfer_to_agent`` on its own. ``HarnessSkillPrefilterPlugin`` asks one ``noul`` question per advertised skill in a single request and rewrites the skill list of that request only. The candidates are judged independently instead of as one ``choice``, because a request often needs two skills and a choice names one winner, and the descriptions travel in the questions so one skill's description cannot decide another skill's answer. The agent instruction keeps every skill, so the next request starts from the full list, and a skill the judgement did not answer for stays advertised: a missing probability is not evidence of irrelevance. ``HarnessAgentRoutingPlugin`` asks one ``choice`` question whose options are the transfer targets and whose option descriptions are theirs, and returns the same ``transfer_to_agent`` call the model would have produced once the judgement clears ``HARNESS_ROUTING_DECISION_THRESHOLD``. Everything less certain stays with the model, which still sees the full transfer instructions, and a judgement never names a target the transfer tool does not expose, so ADK cannot be asked to resolve an agent it does not have. Both are opt-in components (``skill_prefilter`` and ``agent_routing``), judge at most once per invocation, judge only the text of the user's message, and fall back to their rule when no decision model is configured. The READMEs and the environment-variable reference document the five new settings. Change-Id: Ibd735d56a515b50687d48d3592859323d73a6a11
1 parent c57e4ed commit 4b2639b

26 files changed

Lines changed: 2249 additions & 7 deletions

File tree

‎docs/content/docs/references/configuration/environment-variables.en.mdx‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -84,6 +84,11 @@ Prefix `HARNESS_`, used to attach optional Harness plugins to HarnessApp Runtime
8484
| `HARNESS_LONG_RUN_READY_THRESHOLD` | Long-run steering: guidance is injected when the judged probability of being ready is at or above this value; default `0.5`. |
8585
| `HARNESS_MODE_DECISION_THRESHOLD` | Context mode blocks: a block is injected when the judged probability is at or above this value; default `0.5`. |
8686
| `HARNESS_VERIFIER_SUPPORT_THRESHOLD` | Final-answer support: the answer fails when the judged support is below this value; default `0.5`. |
87+
| `HARNESS_SKILL_STRATEGY` | Advertised-skills strategy, `all` or `decision`; default `all`. Requires the `skill_prefilter` component. |
88+
| `HARNESS_SKILL_DECISION_THRESHOLD` | Advertised skills: a skill stays in the request when the judged probability of needing it is at or above this value; default `0.5`. |
89+
| `HARNESS_SKILL_MAX_CANDIDATES` | Advertised skills: a list longer than this is not judged, so every skill stays advertised; default `40`. |
90+
| `HARNESS_ROUTING_STRATEGY` | Sub-agent routing strategy, `model` or `decision`; default `model`. Requires the `agent_routing` component. |
91+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | Routing: the judged agent is transferred to only at or above this probability; default `0.5`. |
8792

8893
The `harness_enhance` block maps to these environment variables when deploying a HarnessApp Runtime. Prefer `harness.yaml` or `veadk agentkit invoke` flags for normal developer workflows; use environment variables for platform integration and container runtimes.
8994

‎docs/content/docs/references/configuration/environment-variables.mdx‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -84,6 +84,11 @@ volcengine:
8484
| `HARNESS_LONG_RUN_READY_THRESHOLD` | 长任务收尾阈值,默认 `0.5`;判定可收尾概率高于该值即注入引导。 |
8585
| `HARNESS_MODE_DECISION_THRESHOLD` | 上下文模式块阈值,默认 `0.5`;判定概率高于该值即注入对应模式块。 |
8686
| `HARNESS_VERIFIER_SUPPORT_THRESHOLD` | 最终回答支撑度阈值,默认 `0.5`;判定支撑度低于该值即判为失败。 |
87+
| `HARNESS_SKILL_STRATEGY` | 技能广告策略,`all` 或 `decision`,默认 `all`;需要 `skill_prefilter` 组件。 |
88+
| `HARNESS_SKILL_DECISION_THRESHOLD` | 技能广告阈值,默认 `0.5`;判定需要该技能的概率不低于该值才继续广告。 |
89+
| `HARNESS_SKILL_MAX_CANDIDATES` | 技能候选上限,默认 `40`;技能数量超过该值时不判定,全部照常广告。 |
90+
| `HARNESS_ROUTING_STRATEGY` | 子 Agent 路由策略,`model` 或 `decision`,默认 `model`;需要 `agent_routing` 组件。 |
91+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | 路由阈值,默认 `0.5`;判定概率不低于该值才直接转移。 |
8792

8893
`harness_enhance` 配置块会在 HarnessApp Runtime 部署时映射为这些环境变量。推荐开发者优先通过 `harness.yaml` 或 `veadk agentkit invoke` 参数启用,环境变量适合平台集成和镜像运行时。
8994

‎docs/extensions/harness/README.md‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -60,6 +60,8 @@ plugins = build_harness_plugins(components=["compactor"])
6060
| `compactor` | `HarnessCompressPlugin` | Compacts oversized tool results and old function responses. |
6161
| `response_verification` | `HarnessResponseVerificationPlugin` | Records tool receipts and checks whether final answers are supported. |
6262
| `long_run_control` | `HarnessLongRunControlPlugin` | Adds finish-oriented guidance when a run approaches its model-call budget. |
63+
| `skill_prefilter` | `HarnessSkillPrefilterPlugin` | Advertises only the skills the current request needs; the agent instruction keeps every skill. |
64+
| `agent_routing` | `HarnessAgentRoutingPlugin` | Transfers to the sub-agent a confident judgement picked; the model routes everything else. |
6365

6466
## Core Concepts
6567

@@ -88,6 +90,10 @@ plugins = build_harness_plugins(components=["compactor"])
8890
| `veadk/extensions/harness/plugins/compactor/` | Tool-result and context compaction callback plugin. |
8991
| `veadk/extensions/harness/plugins/response_verification/` | Receipt recording and final-response verification callback plugin. |
9092
| `veadk/extensions/harness/plugins/long_run_control/` | Long-run guidance callback plugin. |
93+
| `veadk/extensions/harness/modules/skill_prefilter/` | Skill-list parsing and the per-skill judgement. |
94+
| `veadk/extensions/harness/modules/agent_routing/` | Transfer-target judgement. |
95+
| `veadk/extensions/harness/plugins/skill_prefilter/` | Skill-list narrowing callback plugin. |
96+
| `veadk/extensions/harness/plugins/agent_routing/` | Transfer callback plugin. |
9197
| `veadk/extensions/harness/plugins/_shared/` | Internal callback helpers shared by plugins. |
9298
| `veadk/extensions/harness/stores/` | Store protocol and in-memory or JSONL implementations. |
9399

@@ -177,6 +183,8 @@ export HARNESS_VERIFIER_MODE=observe
177183
# export HARNESS_COMPACTION_STRATEGY=decision
178184
# export HARNESS_LONG_RUN_STRATEGY=decision
179185
# export HARNESS_MODE_STRATEGY=decision
186+
# export HARNESS_SKILL_STRATEGY=decision
187+
# export HARNESS_ROUTING_STRATEGY=decision
180188
```
181189

182190
Equivalent YAML:
@@ -220,10 +228,15 @@ veadk agentkit invoke \
220228
| `HARNESS_COMPACTION_KEEP_THRESHOLD` | `0.5` | Compaction candidates: keeps a candidate above this probability. |
221229
| `HARNESS_LONG_RUN_READY_THRESHOLD` | `0.5` | Long-run steering: steers the run to finish above this probability. |
222230
| `HARNESS_MODE_DECISION_THRESHOLD` | `0.5` | Context mode blocks: injects a block above this probability. |
231+
| `HARNESS_SKILL_STRATEGY` | `all` | Advertised skills: `all` or `decision`. |
232+
| `HARNESS_SKILL_DECISION_THRESHOLD` | `0.5` | Skills: a skill stays advertised when its judged probability is at or above this value. |
233+
| `HARNESS_SKILL_MAX_CANDIDATES` | `40` | Skills: a list longer than this is not judged at all, so every skill stays advertised. |
234+
| `HARNESS_ROUTING_STRATEGY` | `model` | Agent routing: `model` or `decision`. |
235+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | `0.5` | Routing: transfers only when the judged agent is at or above this probability. |
223236

224237
## Decision Model Strategies
225238

226-
The four `*_STRATEGY=decision` settings replace a rule with a judgement from
239+
The six `*_STRATEGY=decision` settings replace a rule with a judgement from
227240
the configured decision model. They need `DECISION_MODEL_ENABLED=true` and an
228241
API key; without one, each strategy keeps its rule and logs a warning.
229242

@@ -233,6 +246,8 @@ API key; without one, each strategy keeps its rule and logs a warning.
233246
| `HARNESS_LONG_RUN_STRATEGY` | Model-call counter | Counter, forced after the unconditional count |
234247
| `HARNESS_MODE_STRATEGY` | Precision and artifact keyword markers | Keyword markers |
235248
| `HARNESS_VERIFIER_STRATEGY` | Completion markers plus a successful-receipt check | Builtin rules |
249+
| `HARNESS_SKILL_STRATEGY` | Advertising every loaded skill | Every skill stays advertised |
250+
| `HARNESS_ROUTING_STRATEGY` | The model picking the sub-agent to transfer to | The model routes |
236251

237252
### Judgement Thresholds
238253

@@ -248,6 +263,8 @@ back to `0.5` for an unusable one.
248263
| `HARNESS_LONG_RUN_READY_THRESHOLD` | `0.5` | Steers a run toward its answer sooner |
249264
| `HARNESS_MODE_DECISION_THRESHOLD` | `0.5` | Injects the mode block more often |
250265
| `HARNESS_VERIFIER_SUPPORT_THRESHOLD` | `0.5` | Requires more evidence before the answer passes |
266+
| `HARNESS_SKILL_DECISION_THRESHOLD` | `0.5` | Hides more skills from the list |
267+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | `0.5` | Routes more requests without asking the model |
251268

252269
A judgement also picks an action: long-run steering chooses `narrow_scope` /
253270
`nudge_to_finish` / `force_finish` to shape the injected guidance, and

‎docs/extensions/harness/README.zh.md‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,8 @@ plugins = build_harness_plugins(components=["compactor"])
5555
| `compactor` | `HarnessCompressPlugin` | 压缩过大的工具结果和旧 function response。 |
5656
| `response_verification` | `HarnessResponseVerificationPlugin` | 记录 tool receipt,并检查最终回答是否有证据支撑。 |
5757
| `long_run_control` | `HarnessLongRunControlPlugin` | 当运行接近模型调用预算时,注入面向收敛的引导。 |
58+
| `skill_prefilter` | `HarnessSkillPrefilterPlugin` | 每次请求只广告本次需要的技能;agent 指令本身保留完整列表。 |
59+
| `agent_routing` | `HarnessAgentRoutingPlugin` | 判定足够确信时直接转给对应子 Agent,其余请求仍由对话模型路由。 |
5860

5961
## 核心概念
6062

@@ -83,6 +85,10 @@ plugins = build_harness_plugins(components=["compactor"])
8385
| `veadk/extensions/harness/plugins/compactor/` | 工具结果和上下文压缩回调 plugin。 |
8486
| `veadk/extensions/harness/plugins/response_verification/` | Receipt 记录和最终回答校验回调 plugin。 |
8587
| `veadk/extensions/harness/plugins/long_run_control/` | 长任务收敛引导回调 plugin。 |
88+
| `veadk/extensions/harness/modules/skill_prefilter/` | 技能列表解析与逐候选判定。 |
89+
| `veadk/extensions/harness/modules/agent_routing/` | 转移目标判定。 |
90+
| `veadk/extensions/harness/plugins/skill_prefilter/` | 技能列表收窄回调 plugin。 |
91+
| `veadk/extensions/harness/plugins/agent_routing/` | 转移回调 plugin。 |
8692
| `veadk/extensions/harness/plugins/_shared/` | 多个 plugin 共享的内部回调工具。 |
8793
| `veadk/extensions/harness/stores/` | Store 协议,以及内存 / JSONL 实现。 |
8894

@@ -169,6 +175,8 @@ export HARNESS_VERIFIER_MODE=observe
169175
# export HARNESS_COMPACTION_STRATEGY=decision
170176
# export HARNESS_LONG_RUN_STRATEGY=decision
171177
# export HARNESS_MODE_STRATEGY=decision
178+
# export HARNESS_SKILL_STRATEGY=decision
179+
# export HARNESS_ROUTING_STRATEGY=decision
172180
```
173181

174182
等价 YAML:
@@ -212,17 +220,24 @@ veadk agentkit invoke \
212220
| `HARNESS_COMPACTION_KEEP_THRESHOLD` | `0.5` | 压缩候选:概率高于该值即保留。 |
213221
| `HARNESS_LONG_RUN_READY_THRESHOLD` | `0.5` | 长任务引导:概率高于该值即引导收尾。 |
214222
| `HARNESS_MODE_DECISION_THRESHOLD` | `0.5` | 上下文模式块:概率高于该值即注入。 |
223+
| `HARNESS_SKILL_STRATEGY` | `all` | 技能广告策略:`all` 或 `decision`。 |
224+
| `HARNESS_SKILL_DECISION_THRESHOLD` | `0.5` | 技能:判定概率不低于该值才继续广告。 |
225+
| `HARNESS_SKILL_MAX_CANDIDATES` | `40` | 技能:列表超过该数量时不做判定,全部照常广告。 |
226+
| `HARNESS_ROUTING_STRATEGY` | `model` | 子 Agent 路由策略:`model` 或 `decision`。 |
227+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | `0.5` | 路由:判定概率不低于该值才直接转移。 |
215228

216229
## 判定模型策略
217230

218-
四个 `*_STRATEGY=decision` 开关把一条规则换成判定模型的判定结果,需要 `DECISION_MODEL_ENABLED=true` 与 API Key;没有配置时各自保留原规则并打印告警。
231+
六个 `*_STRATEGY=decision` 开关把一条规则换成判定模型的判定结果,需要 `DECISION_MODEL_ENABLED=true` 与 API Key;没有配置时各自保留原规则并打印告警。
219232

220233
| 策略 | 被替代的规则 | 判定不可用时 |
221234
| --- | --- | --- |
222235
| `HARNESS_COMPACTION_STRATEGY` | 按角色和长度挑选压缩候选 | 内置规则 |
223236
| `HARNESS_LONG_RUN_STRATEGY` | 仅按模型调用次数计数 | 计数规则,超过强制次数后必定生效 |
224237
| `HARNESS_MODE_STRATEGY` | 精度/产物关键词匹配 | 关键词匹配 |
225238
| `HARNESS_VERIFIER_STRATEGY` | 完成类关键词加「有无成功回执」 | 内置规则 |
239+
| `HARNESS_SKILL_STRATEGY` | 广告全部已加载技能 | 技能列表保持不变 |
240+
| `HARNESS_ROUTING_STRATEGY` | 由对话模型选择要转移的子 Agent | 由对话模型路由 |
226241

227242
### 判定阈值
228243

@@ -234,6 +249,8 @@ veadk agentkit invoke \
234249
| `HARNESS_LONG_RUN_READY_THRESHOLD` | `0.5` | 更早把运行推向收尾 |
235250
| `HARNESS_MODE_DECISION_THRESHOLD` | `0.5` | 更频繁注入模式块 |
236251
| `HARNESS_VERIFIER_SUPPORT_THRESHOLD` | `0.5` | 要求更充分的证据才放行回答 |
252+
| `HARNESS_SKILL_DECISION_THRESHOLD` | `0.5` | 从列表里隐藏更多技能 |
253+
| `HARNESS_ROUTING_DECISION_THRESHOLD` | `0.5` | 更多请求不经对话模型直接转移 |
237254

238255
判定还会选动作:长任务引导可选 `narrow_scope` / `nudge_to_finish` / `force_finish` 决定注入的引导文案,最终回答校验可选 `retry_tool_call` / `soften_claim` / `drop_claim` / `ask_user` 决定修复指引;动作不可用时保留默认文案,评级仍然生效。
239256

‎tests/extensions/decisions/test_harness_judges.py‎

Lines changed: 101 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -30,10 +30,17 @@
3030
ToolResultCompactor,
3131
ToolResultCompactorConfig,
3232
)
33+
from veadk.extensions.harness.modules.agent_routing import DecisionAgentRouter
34+
from veadk.extensions.harness.modules.skill_prefilter import DecisionSkillJudge
3335
from veadk.extensions.harness.schemas import CompressionRequest, ConversationMessage
3436

3537
from .fake_system_one import fake_system_one
3638

39+
_ROUTING_AGENTS = {
40+
"billing_agent": "handles invoices and refunds",
41+
"docs_agent": "answers product questions",
42+
}
43+
3744

3845
def _extension(base_url: str) -> DecisionExtension:
3946
return DecisionExtension(
@@ -147,3 +154,97 @@ def test_decision_strategy_end_to_end_keeps_evidence_and_fits() -> None:
147154
assert result.report.compressed_chars <= 12000
148155
assert result.messages[1] == messages[1]
149156
assert result.messages[3] != messages[3]
157+
158+
159+
def _noul_script(
160+
values: dict[str, float],
161+
) -> tuple[int, dict[str, str], dict[str, object]]:
162+
"""Build a response answering ``skill_<index>`` questions."""
163+
return (
164+
200,
165+
{},
166+
{
167+
"model": "fake-system-one",
168+
"answers": {
169+
name: {"type": "noul", "noul": value} for name, value in values.items()
170+
},
171+
},
172+
)
173+
174+
175+
def test_skill_judge_asks_about_every_candidate_in_one_request() -> None:
176+
with fake_system_one([_noul_script({"skill_0": 0.92, "skill_1": 0.08})]) as server:
177+
judge = DecisionSkillJudge(_extension(server.base_url))
178+
probabilities = asyncio.run(
179+
judge.aprobabilities(
180+
user_input="render the chart",
181+
skills={"chart_skill": "draws charts", "mail_skill": "sends mail"},
182+
)
183+
)
184+
185+
assert len(server.calls) == 1
186+
call = server.calls[0]
187+
assert sorted(call.questions) == ["skill_0", "skill_1"]
188+
assert call.questions["skill_1"]["type"] == "noul"
189+
assert probabilities["chart_skill"] == pytest.approx(0.92)
190+
assert probabilities["mail_skill"] == pytest.approx(0.08)
191+
# 候选只出现在问题里:状态没有技能描述,问题之间互相看不见
192+
assert "render the chart" in call.state
193+
assert "draws charts" not in call.state
194+
assert "draws charts" in call.questions["skill_0"]["instructions"]
195+
196+
197+
def test_agent_router_returns_a_target_above_its_threshold() -> None:
198+
scripted = (
199+
200,
200+
{},
201+
{
202+
"model": "fake-system-one",
203+
"answers": {
204+
"target": {
205+
"type": "choice",
206+
"choice": "docs_agent",
207+
"confidence": 0.83,
208+
"probabilities": {"docs_agent": 0.83, "billing_agent": 0.1},
209+
}
210+
},
211+
},
212+
)
213+
with fake_system_one([scripted]) as server:
214+
router = DecisionAgentRouter(
215+
_extension(server.base_url), confidence_threshold=0.8
216+
)
217+
target = asyncio.run(
218+
router.aroute(user_input="how do I rotate a key?", agents=_ROUTING_AGENTS)
219+
)
220+
221+
assert target == "docs_agent"
222+
assert len(server.calls) == 1
223+
call = server.calls[0]
224+
assert call.questions["target"]["criteria"] == _ROUTING_AGENTS
225+
assert "how do I rotate a key?" in call.state
226+
227+
228+
def test_agent_router_leaves_a_low_confidence_choice_to_the_model() -> None:
229+
scripted = (
230+
200,
231+
{},
232+
{
233+
"model": "fake-system-one",
234+
"answers": {
235+
"target": {
236+
"type": "choice",
237+
"choice": "docs_agent",
238+
"confidence": 0.55,
239+
"probabilities": {"docs_agent": 0.55, "billing_agent": 0.45},
240+
}
241+
},
242+
},
243+
)
244+
with fake_system_one([scripted]) as server:
245+
router = DecisionAgentRouter(
246+
_extension(server.base_url), confidence_threshold=0.8
247+
)
248+
target = asyncio.run(router.aroute(user_input="hello", agents=_ROUTING_AGENTS))
249+
250+
assert target is None

0 commit comments

Comments
 (0)