You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
docs: cover the judged verification, recall filter and steering actions
- list the two thresholds the judgement points gained (``HARNESS_VERIFIER_SUPPORT_THRESHOLD``, ``MEMORY_RECALL_RELEVANCE_THRESHOLD``) and say that a threshold compares a rating where the point rates instead of asking
- document ``HARNESS_VERIFIER_STRATEGY`` next to the other strategies, and the
actions the long-run and verification judgements pick
- complete the Harness environment table, which was missing the thresholds the
previous judgement commit added
- document the recall filter in the long-term memory guide
Change-Id: I7c94a9c0c458ccc14133fc29e3369960a7a9d119
|`index`|`str`|`""`| Index/collection name for storing memories. Falls back to `app_name`, then `default_app`. |
32
32
|`app_name`|`str`|`""`| The owning application name; used for data isolation and as the `index` fallback. |
33
33
|`user_id`|`str`|`""`|**Deprecated**, kept only for backward compatibility. |
34
+
|`recall_strategy`|`str`|`"off"`| Recall judgement: with `decision`, the decision model drops memories unrelated to the request. |
35
+
|`recall_relevance_threshold`|`float`|`0.5`| Recall relevance threshold: memories judged below it are not returned. |
34
36
35
37
<Callouttype="info">
36
38
Vector backends (`local`, `opensearch`, `redis`) embed memories, which requires `pip install "veadk-python[extensions]"` and an embedding model (env prefix `MODEL_EMBEDDING_`, falling back to `MODEL_AGENT_API_KEY`). `viking`, `mem0`, `openviking`, and `tos_context` are managed services and need no local embedding.
@@ -183,6 +185,7 @@ The default is `MEMORY_SAVE_STRATEGY=threshold`, the two thresholds above. With
183
185
| Unavailable | Falls back to `MIN_MESSAGES_THRESHOLD` / `MIN_TIME_THRESHOLD`|
184
186
185
187
A skipped turn does not advance the save cursor, so the next accepted judgement writes those events together: nothing is lost. See the [environment variable reference](/references/configuration/environment-variables) for the decision model variables.
188
+
Recall can be judged the same way: with `MEMORY_RECALL_STRATEGY=decision`, `search_memory` rates every returned memory (irrelevant / related but useless / useful background / required) and drops the ones below `MEMORY_RECALL_RELEVANCE_THRESHOLD` (default `0.5`). One judgement covers at most 20 memories, anything beyond that is returned in backend order, and an unavailable judgement keeps every match.
186
189
## Auto-save Memory Policy
187
190
Configure `auto_save_memory_policy` on `Agent` to decide which events are persisted by automatic long-term-memory saving. If omitted, it is equivalent to `"default"`.
| `HARNESS_COMPACTION_KEEP_THRESHOLD` | Compaction candidates: a candidate is kept when the judged probability is at or above this value; default `0.5`. |
84
+
| `HARNESS_LONG_RUN_READY_THRESHOLD` | Long-run steering: guidance is injected when the judged probability of being ready is at or above this value; default `0.5`. |
85
+
| `HARNESS_MODE_DECISION_THRESHOLD` | Context mode blocks: a block is injected when the judged probability is at or above this value; default `0.5`. |
86
+
| `HARNESS_VERIFIER_SUPPORT_THRESHOLD` | Final-answer support: the answer fails when the judged support is below this value; default `0.5`. |
82
87
83
88
The `harness_enhance` block maps to these environment variables when deploying a HarnessApp Runtime. Prefer `harness.yaml` or `veadk agentkit invoke` flags for normal developer workflows; use environment variables for platform integration and container runtimes.
| Long-run steering | `HARNESS_LONG_RUN_READY_THRESHOLD=0.5` | Steers a run toward its answer sooner |
118
120
| Context mode blocks | `HARNESS_MODE_DECISION_THRESHOLD=0.5` | Injects the mode block more often |
121
+
| Final-answer support | `HARNESS_VERIFIER_SUPPORT_THRESHOLD=0.5` | Requires more evidence before the answer passes |
122
+
123
+
Two strategies also choose an action instead of only crossing a threshold, and
124
+
the action shapes what the plugin injects:
125
+
126
+
| Strategy | Action it picks | Effect |
127
+
| --- | --- | --- |
128
+
| Long-run steering | `narrow_scope` / `nudge_to_finish` / `force_finish` | Replaces the injected guidance with the one that fits the trajectory |
129
+
| Final-answer support | `retry_tool_call` / `soften_claim` / `drop_claim` / `ask_user` | Fills the repair instruction the caller hands back to the model |
130
+
131
+
An action that names no known option keeps the default wording; the rating it
132
+
came with is still used.
119
133
120
134
They need a configured decision model; see
121
135
[decisions](../decisions/README.md) for the `DECISION_MODEL_*` variables. A
@@ -125,7 +139,8 @@ Assembling plugins in code selects the same strategies as arguments instead of
0 commit comments