Skip to content

Commit 4dd0ec5

Browse files
committed
fix(memory): compare recall relevance on the threshold's own scale
The recall judge asks a Score question with four levels, so the answer is a probability-weighted position on a 0..3 scale. It was compared straight against ``MEMORY_RECALL_RELEVANCE_THRESHOLD``, which is a ``0..1`` setting: a memory the judgement rated merely "related" (1.0) cleared the default 0.5, and no setting could ask for "required" only, because ``probability_threshold()`` clamps the value to 1.0. The offline tests missed it because their stubs returned probabilities (0.9) rather than the level positions the endpoint actually returns. The position is now scaled onto ``[0, 1]`` before the comparison, so 0.5 means "at least useful" and 1.0 means "required". Checking the plugin against the real TypeSafe endpoint also settled how it is wired to it. The contract holds as documented -- ``POST /v1/systemone`` with Bearer auth, ``state`` / ``model`` / ``questions``, ``noul`` / ``choice`` / ``score`` answers, ``usage``, 401 for a bad key, fail-fast on other 4xx and a retryable set of 429/5xx -- and every judgement point answers, including on CJK state and on eight batched candidates. That check is kept as an opt-in smoke test (``TYPESAFE_RUN_SMOKE=1``) next to the offline ones. Only VeADK's own variable names are read. A key or base URL another tool exported for a provider SDK must not decide the credentials and the endpoint a judgement carrying user text is sent to, so an environment already set up under those names is pointed at the documented one line of bridging instead. Change-Id: Ia2b7bad9b027a3f627a861de6b2a6bdcc349c842
1 parent f52fd30 commit 4dd0ec5

9 files changed

Lines changed: 298 additions & 9 deletions

File tree

‎docs/content/docs/references/configuration/environment-variables.en.mdx‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -123,6 +123,8 @@ model:
123123

124124
`model.decision.*` in `config.yaml` expands to `MODEL_DECISION_*`, which is equivalent to `DECISION_MODEL_*`; when both are set, `DECISION_MODEL_*` wins.
125125

126+
Only these names are read, never a provider SDK's own (such as `TYPESAFE_API_KEY`): a judgement request carries the user's own text, so the credentials and endpoint should not be decided implicitly by a variable another tool exported. An environment already set up under those names needs one line of bridging: `export DECISION_MODEL_API_KEY="$TYPESAFE_API_KEY"`.
127+
126128
When deploying a HarnessApp Runtime, the top-level `decision_model` block in `harness.yaml` maps to this group.
127129

128130
## Long-term memory
@@ -136,7 +138,7 @@ Long-term memory has one optional judgement point for saving and one for recall.
136138
| | `MIN_MESSAGES_THRESHOLD` | Minimum new messages under the `threshold` strategy, default `10`. |
137139
| | `MIN_TIME_THRESHOLD` | Minimum seconds between saves under the `threshold` strategy, default `60`. |
138140
| Recall | `MEMORY_RECALL_STRATEGY` | `off` returns the backend ranking as-is; `decision` rates each memory and drops the low ones; default `off`. |
139-
| | `MEMORY_RECALL_RELEVANCE_THRESHOLD` | Relevance threshold, default `0.5`; memories below it are dropped. Memories the judgement did not rate, and those beyond the 20-item limit, are kept as-is. |
141+
| | `MEMORY_RECALL_RELEVANCE_THRESHOLD` | Relevance threshold, default `0.5`; memories below it are dropped. Relevance is a four-level rating (`irrelevant` / `related` / `useful` / `required`) scaled onto `[0, 1]` (`0` / `0.33` / `0.67` / `1`) before the comparison, so the default keeps memories that are at least `useful`. Memories the judgement did not rate, and those beyond the 20-item limit, are kept as-is. |
140142

141143
## Databases
142144

‎docs/content/docs/references/configuration/environment-variables.mdx‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -123,6 +123,8 @@ model:
123123

124124
`config.yaml` 的 `model.decision.*` 会被展开为 `MODEL_DECISION_*`,与 `DECISION_MODEL_*` 等价;两者同时设置时 `DECISION_MODEL_*` 优先。
125125

126+
只读这套变量名,不读提供商 SDK 自己的变量(如 `TYPESAFE_API_KEY`):判定请求带用户原文,凭证与端点不宜由别人导出的变量隐式决定。已按那些名字配好的环境加一行桥接即可:`export DECISION_MODEL_API_KEY="$TYPESAFE_API_KEY"`。
127+
126128
HarnessApp Runtime 部署时,`harness.yaml` 的顶层 `decision_model` 块映射为这组变量。
127129

128130
## 长记忆
@@ -136,7 +138,7 @@ HarnessApp Runtime 部署时,`harness.yaml` 的顶层 `decision_model` 块映
136138
| | `MIN_MESSAGES_THRESHOLD` | `threshold` 策略下最少新增消息数,默认 `10`。 |
137139
| | `MIN_TIME_THRESHOLD` | `threshold` 策略下两次写入的最小间隔(秒),默认 `60`。 |
138140
| 召回判定 | `MEMORY_RECALL_STRATEGY` | `off` 直接返回后端排序结果;`decision` 逐条评估相关度并丢弃低分项,默认 `off`。 |
139-
| | `MEMORY_RECALL_RELEVANCE_THRESHOLD` | 相关度阈值,默认 `0.5`;低于该值丢弃。未评估到的记忆、以及超出 20 条判定上限的部分原样保留。 |
141+
| | `MEMORY_RECALL_RELEVANCE_THRESHOLD` | 相关度阈值,默认 `0.5`;低于该值丢弃。相关度问的是四档评分(`irrelevant` / `related` / `useful` / `required`),先折算到 `[0, 1]`(`0` / `0.33` / `0.67` / `1`)再比较,默认即「只保留至少 `useful` 的记忆」。未评估到的记忆、以及超出 20 条判定上限的部分原样保留。 |
140142

141143
## 数据库
142144

‎pytest.ini‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -11,3 +11,4 @@ asyncio_mode = strict
1111
markers =
1212
piagent_smoke: real Pi binary/model smoke test (opt in with PIAGENT_RUN_SMOKE=1)
1313
codex_smoke: real Codex binary/sandbox/socket smoke test, stubbed model backend (opt in with CODEX_RUN_SMOKE=1); binds real ports and spawns a subprocess, so it must not run under `pytest -n`
14+
typesafe_smoke: real TypeSafe System One endpoint smoke test (opt in with TYPESAFE_RUN_SMOKE=1 and TYPESAFE_API_KEY); calls a paid, shared service
Lines changed: 222 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,222 @@
1+
# Copyright (c) 2025 Beijing Volcano Engine Technology Co., Ltd. and/or its affiliates.
2+
#
3+
# Licensed under the Apache License, Version 2.0 (the "License");
4+
# you may not use this file except in compliance with the License.
5+
# You may obtain a copy of the License at
6+
#
7+
# http://www.apache.org/licenses/LICENSE-2.0
8+
#
9+
# Unless required by applicable law or agreed to in writing, software
10+
# distributed under the License is distributed on an "AS IS" BASIS,
11+
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12+
# See the License for the specific language governing permissions and
13+
# limitations under the License.
14+
15+
"""Smoke tests against the real TypeSafe System One endpoint.
16+
17+
TypeSafe is a paid, shared service, so these tests are opt in and stay out of a
18+
normal run:
19+
20+
```bash
21+
TYPESAFE_RUN_SMOKE=1 TYPESAFE_API_KEY=... pytest -m typesafe_smoke
22+
```
23+
24+
They answer one question the offline tests cannot: whether the contract the
25+
plugin was written against still holds on the live endpoint.
26+
"""
27+
28+
from __future__ import annotations
29+
30+
import asyncio
31+
import os
32+
33+
import pytest
34+
35+
from veadk.extensions.decisions import (
36+
ChoiceAnswer,
37+
DecisionExtension,
38+
DecisionModelConfig,
39+
NoulAnswer,
40+
ScoreAnswer,
41+
choice_question,
42+
noul_question,
43+
score_question,
44+
)
45+
from veadk.extensions.harness.modules.agent_routing import DecisionAgentRouter
46+
from veadk.extensions.harness.modules.final_response_verifier.support_judge import (
47+
DecisionSupportJudge,
48+
)
49+
from veadk.extensions.harness.modules.invocation_context.mode_judge import (
50+
DecisionModeJudge,
51+
)
52+
from veadk.extensions.harness.modules.long_run_control.judge import (
53+
DecisionConvergenceJudge,
54+
)
55+
from veadk.extensions.harness.modules.skill_prefilter import DecisionSkillJudge
56+
from veadk.extensions.harness.modules.tool_result_compactor import (
57+
DecisionCompactionJudge,
58+
)
59+
from veadk.extensions.harness.schemas import ToolReceipt
60+
from veadk.memory.auto_save_judge import DecisionMemorySaveJudge
61+
from veadk.memory.recall_judge import DecisionRecallJudge
62+
63+
API_KEY = os.environ.get("TYPESAFE_API_KEY", "")
64+
API_BASE = os.environ.get("TYPESAFE_BASE_URL", "https://api.typesafe.ai")
65+
RUN_SMOKE = os.environ.get("TYPESAFE_RUN_SMOKE") == "1"
66+
67+
pytestmark = [
68+
pytest.mark.typesafe_smoke,
69+
pytest.mark.skipif(
70+
not (RUN_SMOKE and API_KEY),
71+
reason="set TYPESAFE_RUN_SMOKE=1 and TYPESAFE_API_KEY to call the live endpoint",
72+
),
73+
]
74+
75+
URGENT_TICKET = (
76+
"Hi, I have been trying to connect my Stripe account for 3 days and the "
77+
"integration keeps failing. I am losing sales. Please help ASAP."
78+
)
79+
80+
81+
def _extension() -> DecisionExtension:
82+
return DecisionExtension(
83+
DecisionModelConfig(enabled=True, api_base=API_BASE, api_key=API_KEY)
84+
)
85+
86+
87+
def test_one_call_answers_all_three_primitives() -> None:
88+
result = _extension().evaluate(
89+
state=URGENT_TICKET,
90+
questions={
91+
"is_urgent": noul_question("Does this message express urgency?"),
92+
"department": choice_question(
93+
"Which team should handle this",
94+
{"billing": "Payment issues", "technical": "Integration problems"},
95+
),
96+
"frustration": score_question(
97+
"How frustrated is the customer?", ["Calm", "Frustrated", "Angry"]
98+
),
99+
},
100+
)
101+
102+
assert result.model, "the response must name the model that answered"
103+
urgent = result.answers["is_urgent"]
104+
assert isinstance(urgent, NoulAnswer) and 0.0 <= urgent.noul <= 1.0
105+
106+
department = result.answers["department"]
107+
assert isinstance(department, ChoiceAnswer)
108+
assert department.choice in {"billing", "technical"}
109+
assert set(department.probabilities) == {"billing", "technical"}
110+
assert sum(department.probabilities.values()) == pytest.approx(1.0, abs=0.01)
111+
assert 0.0 <= department.confidence <= 1.0
112+
113+
frustration = result.answers["frustration"]
114+
assert isinstance(frustration, ScoreAnswer)
115+
assert set(frustration.legend) == {"0", "1", "2"}
116+
assert 0.0 <= frustration.score <= 2.0
117+
assert sum(frustration.probabilities.values()) == pytest.approx(1.0, abs=0.01)
118+
assert 0.0 <= frustration.confidence <= 1.0
119+
120+
121+
def test_environment_configuration_reaches_the_live_endpoint(
122+
monkeypatch: pytest.MonkeyPatch,
123+
) -> None:
124+
"""What ``from_env`` builds must be usable against the real endpoint."""
125+
monkeypatch.setenv("DECISION_MODEL_ENABLED", "true")
126+
monkeypatch.setenv("DECISION_MODEL_API_KEY", API_KEY)
127+
monkeypatch.setenv("DECISION_MODEL_API_BASE", API_BASE)
128+
129+
extension = DecisionExtension(DecisionModelConfig.from_env())
130+
131+
assert extension.enabled
132+
result = asyncio.run(
133+
extension.aevaluate(
134+
state=URGENT_TICKET,
135+
questions={
136+
"is_urgent": noul_question("Does this message express urgency?")
137+
},
138+
)
139+
)
140+
assert isinstance(result.answers["is_urgent"], NoulAnswer)
141+
142+
143+
def test_every_judgement_point_answers_against_the_live_endpoint() -> None:
144+
"""Each judgement point must come back with a usable, in-range answer."""
145+
146+
async def judge_all() -> None:
147+
extension = _extension()
148+
149+
routed = await DecisionAgentRouter(extension).aroute(
150+
user_input=URGENT_TICKET,
151+
agents={
152+
"billing_agent": "handles invoices and refunds",
153+
"docs_agent": "answers product questions",
154+
},
155+
)
156+
assert routed in {"billing_agent", "docs_agent", None}
157+
158+
skills = await DecisionSkillJudge(extension).aprobabilities(
159+
user_input="Draw a sequence diagram for the checkout flow.",
160+
skills={
161+
"archify": "renders architecture and sequence diagrams",
162+
"pptx": "builds slide decks",
163+
},
164+
)
165+
assert set(skills) == {"archify", "pptx"}
166+
assert all(0.0 <= value <= 1.0 for value in skills.values())
167+
168+
kept = await DecisionCompactionJudge(extension).aprotect(
169+
goal="rank the candidates by score",
170+
evidence={
171+
1: "candidate A scored 0.91 with 3 matching skills",
172+
2: "small talk about lunch",
173+
},
174+
)
175+
assert set(kept) == {1, 2}
176+
assert all(0.0 <= value <= 1.0 for value in kept.values())
177+
178+
judgement = await DecisionSupportJudge(extension).areview(
179+
answer="Done, I deployed the service and it is healthy.",
180+
receipts=[
181+
ToolReceipt(
182+
name="run_shell", status="success", summary="wrote notes.md"
183+
)
184+
],
185+
goal="deploy the service",
186+
)
187+
assert judgement.verdict in {"supported", "partial", "unsupported"}
188+
assert 0.0 <= judgement.support <= 1.0
189+
assert 0.0 <= judgement.confidence <= 1.0
190+
191+
modes = await DecisionModeJudge(extension).aprobabilities(
192+
user_input="Refactor the parser and run the tests."
193+
)
194+
assert all(0.0 <= value <= 1.0 for value in modes.values())
195+
196+
convergence = await DecisionConvergenceJudge(extension).ajudge(
197+
goal="make the failing test pass",
198+
trajectory=(
199+
"step 1: ran pytest, 3 failures\n"
200+
"step 2: fixed the fixture, 1 failure left\n"
201+
"step 3: ran pytest again, 1 failure remains"
202+
),
203+
)
204+
assert 0.0 <= convergence.ready <= 1.0
205+
206+
relevance = await DecisionRecallJudge(extension).arelevance(
207+
query="what does the user prefer for diagrams?",
208+
memories=[
209+
"the user prefers diagrams over prose",
210+
"the user lives in Shanghai",
211+
],
212+
)
213+
assert set(relevance) == {0, 1}
214+
assert all(0.0 <= value <= 1.0 for value in relevance.values())
215+
assert relevance[0] > relevance[1]
216+
217+
worth_saving = await DecisionMemorySaveJudge(extension).aworth_saving(
218+
events_text="user: my preferred timezone is Asia/Shanghai; agent: noted."
219+
)
220+
assert 0.0 <= worth_saving <= 1.0
221+
222+
asyncio.run(judge_all())

‎tests/memory/test_memory_recall_judge.py‎

Lines changed: 26 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -38,6 +38,7 @@
3838
)
3939
from veadk.memory.recall_judge import (
4040
MAX_JUDGED_MEMORIES,
41+
MEMORY_RECALL_RELEVANCE_THRESHOLD,
4142
DecisionRecallJudge,
4243
build_recall_judge,
4344
)
@@ -138,8 +139,8 @@ def test_recall_strategy_is_opt_in() -> None:
138139
def test_the_judge_asks_one_score_question_per_memory() -> None:
139140
extension = _StubExtension(
140141
{
141-
"memory_0": ScoreAnswer(score=0.1),
142-
"memory_1": ScoreAnswer(score=0.9),
142+
"memory_0": ScoreAnswer(score=0.0),
143+
"memory_1": ScoreAnswer(score=3.0),
143144
}
144145
)
145146
judge = DecisionRecallJudge(extension) # type: ignore[arg-type]
@@ -148,11 +149,33 @@ def test_the_judge_asks_one_score_question_per_memory() -> None:
148149
judge.arelevance(query="pricing?", memories=["a greeting", "a stated limit"])
149150
)
150151

151-
assert scores == {0: 0.1, 1: 0.9}
152+
assert scores == {0: 0.0, 1: 1.0}
152153
assert set(extension.questions) == {"memory_0", "memory_1"}
153154
assert {question["type"] for question in extension.questions.values()} == {"score"}
154155

155156

157+
def test_the_score_is_scaled_onto_the_threshold_range() -> None:
158+
"""A Score answer arrives on the level scale, the threshold is ``0..1``.
159+
160+
``related`` (level 1 of 4) is below the default ``0.5`` and is dropped;
161+
``useful`` (level 2) is above it and is kept.
162+
"""
163+
extension = _StubExtension(
164+
{
165+
"memory_0": ScoreAnswer(score=1.0),
166+
"memory_1": ScoreAnswer(score=2.0),
167+
"memory_2": ScoreAnswer(score=9.0),
168+
}
169+
)
170+
judge = DecisionRecallJudge(extension) # type: ignore[arg-type]
171+
172+
scores = asyncio.run(judge.arelevance(query="pricing?", memories=["a", "b", "c"]))
173+
174+
assert scores == {0: pytest.approx(1 / 3), 1: pytest.approx(2 / 3), 2: 1.0}
175+
assert MEMORY_RECALL_RELEVANCE_THRESHOLD < 2 / 3
176+
assert MEMORY_RECALL_RELEVANCE_THRESHOLD > 1 / 3
177+
178+
156179
def test_a_memory_without_an_answer_is_rejected() -> None:
157180
extension = _StubExtension({"memory_0": ScoreAnswer(score=0.9)})
158181
judge = DecisionRecallJudge(extension) # type: ignore[arg-type]

‎veadk/extensions/decisions/README.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,16 @@ The same settings can live in `config.yaml`, which VeADK flattens into
5555
`MODEL_DECISION_*` variables. Set both spellings and the explicit
5656
`DECISION_MODEL_*` variable wins.
5757

58+
Only the names above are read, never a provider SDK's own (TypeSafe's
59+
`TYPESAFE_API_KEY`, say): a judgement request carries the user's own text, so
60+
letting a variable exported for another tool decide the credentials and the
61+
endpoint is too implicit. An environment already set up under those names needs
62+
one explicit, auditable line of bridging:
63+
64+
```bash
65+
export DECISION_MODEL_API_KEY="$TYPESAFE_API_KEY"
66+
```
67+
5868
```yaml
5969
model:
6070
agent: {}
@@ -161,6 +171,8 @@ and falls back to the default with a warning for `NaN` or text, which carry no
161171
intent. The thresholds are independent: the same probability costs each point
162172
something different, so raising one must not move the others.
163173

174+
Memory recall asks for a four-level rating (`irrelevant` / `related` / `useful` / `required`), and the answer is scaled onto `[0, 1]` (`0` / `0.33` / `0.67` / `1`) before it is compared with the threshold, so the default `0.5` means "keep only memories that are at least `useful`"; set `1.0` to keep only `required` ones.
175+
164176
A ``noul`` answer is the probability itself and carries no separate confidence,
165177
so its threshold is the whole cascade. An answer that names an option carries
166178
the confidence the model gave that option, and the two points that act on one

‎veadk/extensions/decisions/README.zh.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -49,6 +49,14 @@ DECISION_MODEL_COOLDOWN_SECONDS=30
4949
同样的配置也可以写在 `config.yaml` 里——VeADK 会把配置压平成 `MODEL_DECISION_*`
5050
环境变量。两种写法同时存在时,显式的 `DECISION_MODEL_*` 优先。
5151

52+
只认上面这套变量名,不去读提供商 SDK 自己的变量(如 TypeSafe 的 `TYPESAFE_API_KEY`):
53+
判定请求带着用户原文,凭证和端点由别人导出的变量隐式决定太危险。已经按那些名字配好
54+
的环境,加一行桥接即可,显式且可审计:
55+
56+
```bash
57+
export DECISION_MODEL_API_KEY="$TYPESAFE_API_KEY"
58+
```
59+
5260
```yaml
5361
model:
5462
agent: {}
@@ -147,6 +155,10 @@ agent = Agent(name="router", tools=[decision_evaluate])
147155
或非数字没有原意可保留,回落到默认值并打 warning。阈值之间相互独立——同一个概率
148156
落在不同判定点上代价不同,调高一处不会连带影响其它判定点。
149157

158+
长期记忆召回问的是四档评分(`irrelevant` / `related` / `useful` / `required`),
159+
答案先折算到 `[0, 1]`(`0` / `0.33` / `0.67` / `1`)再与阈值比较,所以默认 `0.5`
160+
的含义是「只保留至少 `useful` 的记忆」;只保留 `required` 就调到 `1.0`。
161+
150162
「是否」类判定返回的概率本身就是级联信号,没有额外的置信度字段,所以它的阈值就是
151163
全部级联。命名选项的判定带着模型给该选项的置信度,两个会据此行动的点可以拒绝没把握
152164
的答案:低于 `HARNESS_VERIFIER_MIN_CONFIDENCE` 时校验回落到内置规则,低于

‎veadk/extensions/decisions/config.py‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -180,7 +180,11 @@ def from_env(cls, env: Mapping[str, str] | None = None) -> DecisionModelConfig:
180180

181181

182182
def _lookup(values: Mapping[str, str], name: str) -> str | None:
183-
"""Read one setting from either accepted environment spelling."""
183+
"""Read one setting from either accepted environment spelling.
184+
185+
Only VeADK's own names are read, never a provider SDK's: a variable another
186+
tool exported must not decide where evidence-bearing judgements are sent.
187+
"""
184188
for prefix in (ENV_PREFIX, CONFIG_YAML_PREFIX):
185189
value = values.get(f"{prefix}{name}")
186190
if value not in (None, ""):

0 commit comments

Comments
 (0)