Add trace-based behavioral tests with Monocle Test Tools - #3
Open
imohammedansari wants to merge 118 commits into
Open
Add trace-based behavioral tests with Monocle Test Tools#3imohammedansari wants to merge 118 commits into
imohammedansari wants to merge 118 commits into
Conversation
…nce#3985) * fix(provisioner): gate legacy skills mount by user visibility * fix(test): aio sandbox provider * fix: use shared legacy skill visibility helper for sandbox mounts
* feat: add composer input polishing * Revert "Merge branch 'main' into feat/input-polish" This reverts commit 5b6cecc, reversing changes made to 45fbc57. * Merge main into feat/input-polish * style(frontend): format input helper polish guard * fix(input-polish): address composer polish review findings Frontend - Add a cancel affordance to the in-flight polish status pill that calls abortInputPolishRequest(), so a slow/hung provider no longer hard-locks the composer for up to stream_chunk_timeout with a page reload (and draft loss) as the only escape. - Reset promptHistoryIndexRef/promptHistoryDraftRef when a rewrite is applied (and on undo), so a stale history-browse index can no longer let the next ArrowDown silently overwrite the polished draft. - Disable polishing while an open human-input card is present, matching the frontend/AGENTS.md rule that composer entry points defer to the card so card-reply metadata is preserved. - canPolishInput now reuses parseGoalCommand/parseCompactCommand instead of a third hardcoded reserved-command regex, and drops the phantom /help entry (no /help parser exists in the composer), so future builtins only need to be taught to the existing parsers. Backend - Extract the non-graph one-shot LLM path (build model + inject Langfuse metadata + system/user invoke + text extract) into deerflow.utils.oneshot_llm.run_oneshot_llm, shared by the input-polish and suggestions routers so tracing-metadata and invocation shape cannot drift between the two copies. - strip_think_blocks gains truncate_unclosed (default True, preserving the suggestions/goal JSON-prep behavior); input polish passes False so a draft that legitimately contains a literal <think> substring is no longer truncated into a partial rewrite or a spurious 503. - Validate the empty-check and max_chars boundary against the same stripped view of the draft that is sent to the model, so the user-facing length boundary and the model input can no longer disagree. Tests / docs - Backend: literal-<think> preservation, whitespace-only rejection, and normalized-length/model-input agreement cases; suggestions tests repoint the create_chat_model patch to the shared helper module. - Frontend: helper unit tests updated for the /help/reserved-command change; a new Playwright case covers cancelling an in-flight polish request. - backend/AGENTS.md documents the shared one-shot helper and the polish normalization/think-tag behavior. --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
…ance#3981) * feat(frontend): render slash-skill activations as inline chips Show an explicit `/skill` activation as a compact inline chip in both the composer and the chat transcript instead of raw slash text. - Composer: selecting a skill suggestion stores it as a removable chip aligned inline with the textarea; the leading `/skill ` prefix is reattached only at submit time, so the backend activation protocol is unchanged. Backspace on an empty input or the chip's close button clears it; history navigation is disabled while a chip is active. - Transcript: human messages that begin with `/skill` render the skill as a read-only chip followed by the task text. - Add a shared `core/skills/slash.ts` (`parseSlashSkillReference` + `resolveSlashSkillDisplay`) mirroring the backend `slash.py` gate, so the transcript only shows a chip when the skill actually exists and is enabled. This removes a duplicated regex/reserved-name list and keeps display semantics consistent with backend activation. Add unit tests for the shared slash parser and extend the chat e2e to assert the composer still submits `/skill <task>` after showing a chip. * chore(frontend): format chat e2e test * refactor(skills): address slash-skill chip review feedback Follow-up to the inline slash-skill chip PR, resolving three second-order review findings: - Drive the reserved-command set and skill-name grammar from a shared contracts/slash_skill_contract.json instead of a hand-copied "keep in sync" pair. slash.ts and slash.py now reference the fixture, and contract tests on both sides fail CI if either drifts. - Extract a shared SlashSkillChip so the composer and transcript chips stay in lockstep, and normalize the off-scale /8 and /12 opacity steps to the standard /10 and /20 tokens. - Split HumanMessageText into a pure parse gate plus a slash-only subtree that owns the useSkills() lookup, so a skill-enabled toggle no longer re-renders every plain-text human turn. Verified: frontend eslint + tsc clean, pnpm test 572 pass (incl. new slash-contract test); backend slash contract + slash-skills tests 31 pass. * style(tests): sort slash skill contract imports * fix(composer): inline the slash-skill text so the chip aligns with input Address the "composer body layout change" review on bytedance#3981 by rendering the active skill as an inline chip in the same text flow as the prompt, rather than a separate flex row that drifted the box model across states. - Render the chip + prompt inside one leading-6 wrapper and edit the prompt through a `contentEditable` span, so the chip sits inline with the first line and long/multi-line input wraps naturally back to the container edge. - Align the chip with `align-top`: its h-6 (24px) height matches the text line height, so chip and first-line centers coincide exactly (measured delta 0), fixing the chip being raised above the baseline. - Restore the placeholder in chip mode via a `data-empty` CSS `::before`, which also gives the empty editable span width so it is no longer treated as hidden. - Widen the IME helper to `HTMLElement` and route the span's keydown/paste through the shared skill-suggestion, prompt-history, backspace-to-clear, and IME-composition handlers so contentEditable behaves like the textarea. - Extend chat.spec.ts to drive the inline skill editor instead of the textarea after a chip is shown. * style(frontend): fix composer class order formatting * fix(composer): break long unbroken input inside the slash-skill row The inline slash-skill editor wrapped with `break-words` (overflow-wrap: break-word), which only moves an over-long token to the next line before breaking it. A long unbroken string therefore started on the line below the chip, and when the string contained a break opportunity such as a hyphen the browser wrapped there and pushed the remaining run to the next line, leaving a wide gap on the right. Switch to `break-all` (word-break: break-all) so the text fills each line from the chip and packs tightly regardless of hyphens or CJK.
…ken_budget (bytedance#3875 Phase 2) (bytedance#3980) Phase 2 of bytedance#3875. Two guardrail axes can end a subagent run early — the turn budget (GraphRecursionError) and the token budget (TokenBudgetMiddleware) — and both now surface *why* through one additive `subagent_stop_reason` field instead of a status enum. This completes and course-corrects Phase 1 (bytedance#3949), which shipped the turn-budget cap as a `max_turns_reached` status enum. The agreed Phase 2 design replaces that enum with an optional `stop_reason` field (token_capped | turn_capped | loop_capped): a new enum value would break v1 consumers, while an additive field is ignored by older frontends and ledger readers. `max_turns_reached` and SubagentStatus.MAX_TURNS_REACHED are removed. - subagents.token_budget config (default enabled, 2,000,000 tokens, warn 0.7) with per-agent override; TokenBudgetMiddleware is now attached in build_subagent_runtime_middlewares so the cost-ceiling backstop engages for every subagent. The hard-stop does not raise — it strips tool_calls and lets the run finish with a final answer, recording the cap on a per-run consume_stop_reason() accessor. - executor.py: on normal completion it reads consume_stop_reason() and stamps completed + token_capped when the budget fired; on GraphRecursionError it recovers the last AIMessage partial (completed + turn_capped) or, if nothing usable survived, failed + turn_capped. SubagentResult gains stop_reason. - status_contract.py / contracts/subagent_status_contract.json (v2) / frontend subtask-result.ts: additive subagent_stop_reason field, pinned by test_status_values_match_contract / test_stop_reason_values_match_contract. - task_tool.py + delegation_ledger.py: drop the max_turns_reached paths; the ledger captures stop_reason and renders model-facing "capped" guidance so the lead reuses a capped completion knowingly. The 2,000,000-token default is deliberately loose (tighten to taste) — it would have roughly halved the reported 4.4M burn while leaving legitimate deep-research runs (max_turns=150) room. Subagent summarization is a follow-up.
bytedance#4002) * fix(security): neutralize prompt-injection tags in remote tool results User input is already neutralized for framework/injection tags, but tool results are not. Remote content fetched by web_fetch/web_search is equally untrusted and can carry a forged <system-reminder> block that reaches the model verbatim as authoritative context. Extract a shared neutralize_untrusted_tags() primitive from InputSanitizationMiddleware and apply it to remote-content tool results (web_fetch/web_search/image_search) via a new ToolResultSanitizationMiddleware. Local tool output (bash/read_file) is left untouched so legitimate code/file content is never mangled. * test: update subagent middleware count for tool-result sanitizer The new ToolResultSanitizationMiddleware adds one entry to the shared runtime chain (11 -> 12). Update the subagent count assertion, use a lazy import for neutralize_untrusted_tags so the module loads even when tests stub the input-sanitization module, and document the new middleware in AGENTS.md. * fix(security): address review — sanitize bare str list items; document MCP scope - Neutralize bare str elements inside a ToolMessage content list (previously only {type:text} dict blocks were rewritten), matching the str-in-list shape ToolOutputBudgetMiddleware._message_text already anticipates. - Document the name-based allowlist limitation: MCP remote-content tools registered under arbitrary names (e.g. fetch_url) are not covered; a name heuristic is avoided to prevent mangling local tool output, with metadata tagging tracked as a follow-up. Add a regression test pinning this boundary.
…ance#3987) * feat(helm): add production-ready Helm chart for Kubernetes deployment Adds deploy/helm/deer-flow, a native-Kubernetes translation of the production docker-compose stack, plus CI to publish its images and chart. * ci(release): gate releases on version-source consistency Add a reusable verify-versions workflow invoked by both chart.yaml and container.yaml on v* tags. It runs scripts/verify_versions.sh against the tag and fails the release — skipping all image and chart publishing — when Chart.yaml (version + appVersion), backend/pyproject.toml, or frontend/package.json don't all match the tag. Add scripts/verify_versions.sh (the check, also runnable locally) and scripts/bump_version.sh (bumps all four sources in lockstep, then self-verifies). Document the release flow in RELEASING.md and link it from AGENTS.md. * fix(deploy): address Helm chart review feedback (bytedance#3987) Three review items from willem-bd: 1. nginx IPv6 listen strip never matched. The sed pattern required a `;` immediately after `2026`, but the rendered config emits `listen [::]:2026 default_server;` (space + `default_server` before the `;`), so the line was never deleted and nginx crash-looped on pods without IPv6 (`socket() :::2026 failed (97: Address family not supported)`). Drop the trailing `;` from the pattern so it matches. Same latent bug fixed in docker-compose-dev.yaml. 2. Passwords were spliced into DSNs verbatim, so a password containing URL-special chars (@ : / # ? % [ ] space) produced a malformed DSN and a confusing parse error. Add a `deer-flow.urlEscape` helper (replace-based: Sprig lacks urlqueryescape, and regexReplaceAllLiteral treats the replacement as a regex template so `[`/`]`/`?` break it) and apply it to the password in the postgres and redis DSNs. The raw `postgres-password` / `redis-password` keys stay unencoded - they back POSTGRES_PASSWORD / REDIS_PASSWORD, not a URL segment. 3. NODE_HOST defaulted to "gateway", which can never route: the gateway Service is ClusterIP:8001 and knows nothing of a sandbox NodePort, so a user who skips the caveat gets unreachable sandboxes with no error at install time. Default NODE_HOST to the provisioner pod's node IP via the downward API (status.hostIP) - a NodePort is exposed on every node, so <node-IP>:<NodePort> routes from the gateway on most clusters. `provisioner.nodeHost` remains an override for CNIs/policies that block pod->node-IP traffic. Updated NOTES.txt, values.yaml, and the chart README. (bytedance#3929 remains the long-term fix - ClusterIP + cluster-DNS URL removes NODE_HOST and the NodePort exposure entirely.) Validated with helm lint, helm template (incl. a special-char password rendering the encoded DSNs), and a sed pattern-match check. * fix(deploy): address round-2 Helm chart review feedback (bytedance#3987) Three "Medium" items from willem-bd: 1. No helm lint / helm template gate before publish. A template regression ships as an immutable OCI artifact (GHCR won't overwrite --version), so gate packaging on `helm lint` + `helm template --include-crds` in chart.yaml before `helm package`. (ct lint / helm-unittest deferred.) 2. Action pinning inconsistent + PR body overstates it. SHA-pin actions/checkout (v6.0.3, df4cb1c0) and actions/attest-build-provenance (v2.4.0, e8998f94) across the publishing workflows (chart.yaml, container.yaml, verify-versions.yml), matching the existing docker/* SHA-pin pattern. Resolves the checkout @v4/@v6 mismatch and makes the "SHA-pinned actions" claim accurate. Other pre-existing workflows left untouched (out of scope for this PR). 3. Provisioner RBAC broader than needed. Dropped the unused update/patch verbs and the pods/exec + events rules from the provisioner Role - audited against docker/provisioner/app.py, which only calls get/create/delete on pods and get/list/create/delete on services. Fixed NOTES.txt to accurately describe the grant instead of understating it as "create Pods and Services". The remaining scope concern - verbs apply to all Pods in the namespace, not just sandbox Pods - is still deferred (RBAC can't scope by label; needs a dedicated namespace or admission control), now noted in NOTES.txt and README. Validated with helm lint + helm template (narrowed Role renders with exactly get/list/watch/create/delete). * feat(helm): enable sandbox+web tools out of the box The chart's default config loaded zero agent tools (config.tools empty -> "Total tools loaded: 0"), so a fresh install gave an agent that could do nothing useful. Add tool_groups + tools to the default config block: - web: web_search (ddg), web_fetch (jina), image_search - no API key - file:read: ls, read_file, glob, grep - file:write: write_file, str_replace - bash The file/bash tools run inside the AIO sandbox the chart already configures; the web tools need outbound internet from the gateway pod (swap backends or drop entries for air-gapped clusters - see config.example.yaml). Also bump config_version 15 -> 19 to match config.example.yaml (the chart had drifted behind). NOTES.txt and the README example updated to match. * ci(helm): add chart validation + config_version drift check on PR Extend the chart workflow with a PR-triggered validate-chart job that runs helm lint, helm template --include-crds, and a config_version drift check: it parses config_version from both config.example.yaml and the chart's values.yaml and fails the build (with a ::error:: naming the files to bump) if the chart is behind the example. This catches the kind of drift this PR is fixing - the chart sat at v15 while the example moved to v19 - before it can merge again. verify-versions and publish-chart stay tag-only; publish-chart now needs: [verify-versions, validate-chart]. validate-chart runs on both PRs and tag pushes: the tag arm is required because a job that `needs` a skipped job is itself skipped under the default success() check, so validate-chart must actually run on tag pushes or publish-chart would never fire. * Bump config version to 20
* feat: add MCP routing hints * test: isolate mcp routing prompt config * fix: address mcp routing review feedback
* fix: recover from empty tool call names * test: harden empty tool call recovery --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
…edance#4015) The tag-publish workflow (container.yaml) built backend/Dockerfile with no build-args, so UV_EXTRAS was empty and the published *-backend image shipped without the Postgres driver (only --extra redis). Multi-replica deployments (K8s/Helm) that need shared Postgres persistence instead of file-based SQLite could not use the release image without rebuilding it. Pass UV_EXTRAS=postgres so the release image includes deerflow-harness[postgres]. Additive only: single-replica sqlite/redis setups keep working; the Postgres driver is added, mirroring how redis is already always baked in.
…ce#4013) Bumps [langsmith](https://github.com/langchain-ai/langsmith-sdk) from 0.8.0 to 0.8.18. - [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases) - [Commits](langchain-ai/langsmith-sdk@v0.8.0...v0.8.18) --- updated-dependencies: - dependency-name: langsmith dependency-version: 0.8.18 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ce#4001) * bench: add provider-agnostic sandbox benchmark with BoxLite warm pool results - scripts/bench_sandbox_provider.py: CLI for measuring acquire/run/release across providers, scenarios, workloads, and concurrency levels - scripts/summarize_bench.py: JSONL aggregation with p50/p95/p99 tables - bench_results.jsonl: 110 turns across 7 scenarios on real BoxLite 0.9.7 Key findings: cold acquire: ~860ms warm reclaim: ~14ms (60x speedup) release: ~0ms warm_hit_rate: 95% (warm_same_thread) * perf(boxlite): skip health check for recently-released warm pool boxes Boxes released within health_check_skip_seconds (default 5.0s) are promoted directly without the ~14ms echo-ok round-trip. A VM alive seconds ago is overwhelmingly likely to still be alive. Add sandbox.health_check_skip_seconds config option. Set to 0 to always health-check (old behaviour). Benchmark (warm_same_thread, noop, 20 iters): acquire p50: 14.9ms → 0.0ms total p50: 29.9ms → 14.0ms * chore: move benchmark scripts into backend/scripts/benchmark/ * fix: address BoxLite benchmark review findings * fix(boxlite): only skip warm reclaim checks for released boxes * fix(benchmark): keep BoxLite shim workaround off the event loop * fix(boxlite): invalidate dead boxes from command path * test(boxlite): cover skip window and invalidation edge cases * fix(boxlite): treat sandbox-has-been-closed as terminal in _exec * fix(boxlite): harden warm-pool reclaim and benchmark accounting * fix(boxlite): validate warm-pool reclaims by default * fix(config): expose boxlite health-check skip setting * fix(boxlite): tighten failure classification and benchmark workaround * Update config_version to 21 in values.yaml --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
…bytedance#4014) Bumps [pydantic-settings](https://github.com/pydantic/pydantic-settings) from 2.14.0 to 2.14.2. - [Release notes](https://github.com/pydantic/pydantic-settings/releases) - [Commits](pydantic/pydantic-settings@v2.14.0...v2.14.2) --- updated-dependencies: - dependency-name: pydantic-settings dependency-version: 2.14.2 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…ounts (bytedance#4016) * feat(provisioner): add ClusterIP mode for sandbox services * feat(provisioner): support optional skills pvc subpath
imohammedansari
force-pushed
the
monocle-test-tools
branch
from
July 9, 2026 22:20
b9db1cc to
8a1f22a
Compare
Trace-based tests under tests/monocle/ asserting against the agent's Monocle execution traces: 4 offline tests load a recorded trace by file (with_trace_source), one per curated question, plus 1 live end-to-end test. Fluent structural asserts (agent, tools, input/output, token/duration budget); additive only, no app-code changes.
…twork sinks (bytedance#4130) python-env-dump-exfil flags a file that both reads the bulk process environment and reaches a network sink. The call-based sink check only listed requests get/post/put/request and httpx get/post, so a bulk env dump sent through an equally body-carrying method (requests.patch/delete, httpx.put/patch/delete, or the generic httpx.request/stream) evaded the CRITICAL finding whenever the destination URL was not a plain string literal (e.g. built at runtime) -- the exact evasion the string-literal URL sink is meant to resist. requests.post was caught but requests.patch was not, an arbitrary gap on clients the analyzer already covers. Complete the requests and httpx HTTP-verb surface in _call_is_network_sink so an obfuscated-URL exfil through those methods is flagged like post/put. Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
…er (bytedance#4129) _channel_storage_user_id is the single source of truth for a channel run's identity: _resolve_run_params resolves the owner into run_context["user_id"]. The /skill whitelist pre-check for a per-user custom agent (_resolve_available_skill_names) resolves that owner and then drops it, calling load_agent_config(name) with no user_id. It falls back to get_effective_user_id(), but this runs on the ChannelManager dispatch loop where the _current_user contextvar is never set, so it resolves "default" -- reading users/default/agents/{name}/ instead of the owner's bucket. When that bucket has no such agent (the common case) load_agent_config raises FileNotFoundError, which the dispatch loop turns into "An internal error occurred" on every /skill command; when a foreign agent shares the name, the whitelist is decided by the wrong user's skills list. Pass the resolved owner (run_context["user_id"]) to load_agent_config, matching every other caller (gateway/routers/agents.py, update_agent_tool.py, github/registry.py). None when no owner is resolvable, preserving the prior default-user behavior for unbound/no-auth channels.
bytedance#4086) The select: query form lets the model explicitly name the tools it wants to promote. Capping the result at MAX_RESULTS (5) silently drops valid selections when more than 5 deferred tools are requested at once. This removes the [:MAX_RESULTS] slice on the select: branch in DeferredToolCatalog.search() (exact-by-name lookups should return every matched tool) and the redundant re-cap in build_tool_search_tool() (search() already caps non-select query forms internally). Co-authored-by: Claude <noreply@anthropic.com>
…wner (bytedance#4108) * refactor(sandbox): give the host→virtual output-mask regex a single owner Two call sites rewrite host paths back to their virtual form in text that reaches the model — LocalSandbox._reverse_output_patterns (bash output) and sandbox.tools._compiled_mask_patterns (glob/grep/ls results) — and each built the same `escape(base) + boundary + tail` rule from its own copy. That duplication has already produced two bugs: bytedance#4035 added the segment boundary to the reverse patterns and missed the masking patterns, and bytedance#4053 had to add the same boundary to the other copy. Extract the rule into sandbox/path_patterns.py so a third copy cannot silently disagree. The extraction is not a pure move: the two sites disagree on the base. tools.py derives bases from _path_variants (which yields Windows spellings) and matches them against output whose separators it does not control, so it relaxes the separators inside the base; LocalSandbox resolves its bases from the running platform and must not be widened. That difference is now an explicit `separator_agnostic` parameter rather than an accident of two implementations. The boundary and tail constants are private: build_output_mask_pattern is the only supported spelling, so a third site cannot import the pieces and hand-roll a variant. Behavior is unchanged at both sites — pinned by tests that reproduce each pre-extraction expression byte-for-byte. * test(sandbox): pin the base the helper must not normalize Review notes on bytedance#4108. The committed snapshot compares the helper against hand-copied literals of the pre-extraction expressions, so its red-ness rests on those literals, not on the length of _BASES -- both sides compute the same expression, and 5k fuzzed bases produce zero byte-differences. Mutating the helper one clause at a time (12 mutations over the boundary, the tail and the escape/replace) shows the seven committed bases catch 11: the miss is a helper that normalizes its input by rstripping a trailing separator. Only a trailing-slash base or a Windows drive root catches that, and Path.resolve() / str(Path(...)) strip trailing slashes, so neither call site can produce the former. C:\ survives resolve() with its separator intact, so that is the one base worth adding. Also point local_sandbox's comment at path_patterns, the owner, instead of citing _content_pattern as the class reference, and drop the rationale it now duplicates from the owner's docstring -- a second copy of the explanation drifts the same way the second copy of the regex did. The site-specific half stays. Comments and test data only; no behavior change.
…ytedance#4103) * fix(skills): activate a slash skill once per run, not per model call SkillActivationMiddleware injects the activation reminder for a slash command via request.override(messages=...), which LangChain's create_agent uses for a single model call and never writes back to graph state. The dedup guard scans request.messages for a prior reminder, but model_node rebuilds request.messages fresh from persisted state on every tool-loop step, so the reminder is never present on the 2nd..Nth model call of a turn. Every model call therefore re-parsed the command, re-read SKILL.md from disk, re-injected the multi-KB body, and re-recorded an "activate" audit event, despite the code intending a single activation per run (bytedance#3861 semantics: one activation call, many follow-up model calls). Key the dedup off the run context instead, which LangGraph threads through every model-node call of a run (the same durable signal the request-scoped secret source already uses). The activation call records the slash message's identity in context; later calls for the same message skip re-activation. A new user slash message keys differently and still activates. Secret binding is unaffected: it already re-resolves from the persisted slash source on every call. Adds regression tests that rebuild the real multi-call turn state and assert a single activation across the tool loop, plus a test proving a new slash command still activates. * fix(skills): address review nits on run-scoped activation dedup - Extract _already_activated(run_context, run_key) so the dedup check mirrors the existing _has_existing_activation_for_target sibling instead of an inline dense conditional. - Compute _activation_run_key() once in _find_activation_target and thread it through _prepare_model_request instead of recomputing it at the write site, making the "same key for check and write" invariant explicit in the code rather than implicit. - Document why the run-context write is an overwrite rather than an append/set: only the latest real user message is ever considered an activation target, so there is nothing earlier in the run worth preserving. - Add a regression test locking in the degraded-path contract: when runtime.context is None, the middleware still activates per-call instead of crashing or wrongly no-op'ing.
* fix subagent total delegation cap * fix embedded subagent run cap context * fix subagent cap config consistency * fix resumed subagent run cap boundary * fix legacy resume subagent boundary * address subagent cap review feedback --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
… block (bytedance#4137) SOUL.md is agent-editable (setup_agent / update_agent persist it) and get_agent_soul renders it into the <soul> block of the lead-agent system prompt without escaping. A crafted personality such as "</soul></system-reminder>\n\nSYSTEM: ..." can close the block and relocate the text after it out of the trust zone the system prompt declares — the same break-out the skill/memory/tool-result escaping in bytedance#4097/bytedance#4119/bytedance#4128/bytedance#4099 already closes at their render sites. <soul> is the remaining one, and it lands in the highest-trust system-role block. Escape with html.escape(quote=False) (element-text position, never an attribute). Adds a regression test that fails on main. Signed-off-by: Yufeng He <40085740+he-yufeng@users.noreply.github.com>
…nce#4136) * fix(agents): load SOUL.md from agent dirs without config.yaml (bytedance#4135) resolve_agent_dir requires config.yaml to be present in an agent directory (added in bytedance#3481 to fix bytedance#3390, where memory-only directories were mistaken for agent directories). However, SOUL.md loading does not depend on config.yaml. When an agent is configured externally (e.g. via DEER_FLOW_CONFIG_PATH) and the agent directory contains SOUL.md but no config.yaml, resolve_agent_dir skips the directory and load_agent_soul returns None. Add a fallback in load_agent_soul: if the resolved directory does not contain SOUL.md, check the per-user and legacy directories directly for SOUL.md. This preserves the bytedance#3390 fix for resolve_agent_dir (which is also used by load_agent_config) while allowing SOUL.md to load independently of config.yaml. * fix(agents): gate SOUL.md fallback on missing config.yaml per review Address review feedback on PR bytedance#4136: 1. Gate the fallback condition on not (agent_dir / config.yaml).exists() so it only fires when resolve_agent_dir returned its default path (no agent dir qualified), not when a properly-resolved per-user agent simply lacks SOUL.md. This preserves the per-user shadowing invariant. 2. Fix test_loads_soul_from_user_dir_without_config_yaml to actually exercise the fallback path: per-user memory-only dir (no config.yaml, no SOUL.md) + legacy dir with SOUL.md (no config.yaml) -> fallback finds legacy SOUL.md. 3. Add test_soul_not_leaked_from_legacy_when_per_user_has_config to verify the gate prevents legacy SOUL.md leaking into a per-user agent. 4. Patch get_effective_user_id in test_loads_soul_without_config_yaml for consistency and resilience.
… subclasses (bytedance#4102) * fix(models): apply stream_chunk_timeout default to all BaseChatOpenAI subclasses The 240s stream_chunk_timeout default (issue bytedance#3189, PR bytedance#3195) was scoped to a class-path allowlist of only ChatOpenAI and PatchedChatOpenAI. Every other OpenAI-compatible provider that subclasses BaseChatOpenAI — VllmChatModel, MindIEChatModel, PatchedChatDeepSeek, PatchedChatMiMo, PatchedChatStepFun and PatchedChatMiniMax — was excluded, so they kept langchain-openai's aggressive 120s built-in chunk-gap timeout and, worse, silently discarded a user's explicit stream_chunk_timeout override from config.yaml. Issue bytedance#3189 was itself reported on mimo-v2.5 (PatchedChatMiMo), the exact class the original fix left out. Gate the injection on issubclass(model_class, BaseChatOpenAI) instead of the string allowlist, so any OpenAI-compatible subclass inherits the default and honors an explicit override. Genuinely non-OpenAI clients (e.g. ChatAnthropic) stay excluded and still have the kwarg dropped before it reaches a constructor that would divert it into model_kwargs and fail at request time. * fix(models): address review nits on stream_chunk_timeout default Correct the module-level comment above _DEFAULT_STREAM_CHUNK_TIMEOUT_SECONDS: langchain-openai's built-in stream_chunk_timeout default is 120s, not 60s (BaseChatOpenAI.stream_chunk_timeout's default_factory reads LANGCHAIN_OPENAI_STREAM_CHUNK_TIMEOUT_S with a 120.0 fallback). Simplify the BaseChatOpenAI gate in _apply_stream_chunk_timeout_default from `isinstance(model_class, type) and issubclass(model_class, BaseChatOpenAI)` to just `issubclass(...)`. The sole caller passes model_class from resolve_class(), which already raises before returning anything that isn't a type, so the isinstance half can never be False there. Also soften the docstring's non-OpenAI-client bullet: ChatAnthropic declares extra="ignore" and silently drops an unrecognized kwarg rather than diverting it into model_kwargs and failing at request time (that failure mode is specific to other OpenAI-style clients).
…nce#4064) * fix(runs): cancel degrades to lease takeover for multi-worker Work item 4 of the multi-worker ownership epic (bytedance#3948). Problem: POST /runs/{run_id}/cancel landing on a non-owning worker returns 409 — the cancel button silently fails under GATEWAY_WORKERS>1 with no sticky routing. cancel() required the current worker to hold the in-memory task/abort_event, which any non-owner pod cannot satisfy. Changes: - RunManager.cancel() returns CancelOutcome enum (cancelled / taken_over / lease_valid_elsewhere / not_active_locally / not_cancellable / unknown) instead of bool, so the router can map each outcome to the right HTTP response. - New store primitive claim_for_takeover(): a single atomic conditional UPDATE that marks a run as error only when status IN (pending, running) AND (lease IS NULL OR lease < now - grace). Closes the stale-read / concurrent-heartbeat race — if the owner renews between our read and write, the UPDATE matches 0 rows and we surface lease_valid_elsewhere. - HTTP cancel + stream-join endpoints route on CancelOutcome: cancelled -> 202 (or 204 with wait=true); taken_over -> 202 immediately (no SSE streaming — the run is terminal on another worker, streaming would hang); lease_valid_elsewhere -> 409 + Retry-After header computed from lease_expires_at + grace_seconds. - RunManager.grace_seconds exposed as a public property; the router no longer reaches into _run_ownership_config. - _is_lease_expired extracted to a module-level function, shared by RunManager.cancel() and MemoryRunStore.claim_for_takeover(). - GATEWAY_WORKERS=1 + heartbeat_enabled=false is zero-regression: the non-local path short-circuits to not_active_locally, preserving the original 409 behaviour the existing tests pin. Tests: 12 new (5 store primitive + 4 cancel-takeover unit + 3 HTTP including a regression guard verifying POST /stream?action=interrupt on a dead-owner run returns 202 instead of hanging on SSE). 244 directly-related tests pass; 36/36 blocking-IO gate pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(runs): guard update_status and self-terminate on takeover Two defenses close a split-brain window where the original owner could overwrite a peer's takeover status: - update_status (SQL + memory store) now guards on status IN ('pending','running'). When takeover already set the row to 'error', the owner's final status write matches 0 rows and is dropped. - _persist_status: when update_status returns False, check whether the row exists before attempting recovery via put(). If the row exists (takeover by another worker), skip recovery instead of blindly upserting over the takeover. - Heartbeat _renew_leases: when update_lease returns False (row no longer pending/running or owner changed), cancel the local task so wasted CPU is bounded to the next heartbeat tick (~10s) instead of the full task lifetime. Also fix three reviewer feedback items: - Re-fetch the store row when cancel() returns lease_valid_elsewhere, so Retry-After uses the owner's freshly-renewed lease instead of a stale value from request start. - Fallback 'unknown' in takeover error message when owner_worker_id is NULL (pre-ownership data). - Remove dead else-10 branch from grace_seconds property (unreachable — all callers are downstream of the heartbeat_enabled guard). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(runs): pin split-brain defences from update_status guard + heartbeat Three tests lock down the takeover authoritativeness so a late-running owner cannot overwrite a peer's claim: - update_status must reject writes when the store row is already terminal (taken over by another worker). - _persist_status must skip row-recovery via put() when the row exists but has been taken over. - Heartbeat _renew_leases must cancel the local task when update_lease returns False (row claimed by another worker). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(runs): precise outcome + log when local cancel loses to peer takeover Two reviewer precision nits on the split-brain defence: - _persist_status: branch the skip-reason log on existing["status"]. error → WARNING "peer takeover" (anomalous); interrupted/success → INFO "local cancel/completion race" (expected when user hits stop as the run finishes). Stops noisy false-positive takeover warnings in operator logs. - cancel() local path: when _persist_status returns False, re-check the store. If a peer's claim_for_takeover flipped the row to error between our in-memory cancel and the guarded update_status, surface taken_over instead of cancelled so the client sees a status consistent with the store. Test: test_cancel_returns_taken_over_when_peer_claims_during_local_cancel pins the race outcome. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(runs): widen update_status guard, de-duplicate lease helpers, add coverage Round 3 of reviewer feedback: - Widen update_status guard to status IN ('pending','running','interrupted'). The original guard blocked interrupted→error (the rollback finalize path), losing the "Rolled back by user" message. interrupted is now permitted while error/success stay locked — takeover protection unchanged. - claim_for_takeover False now re-reads the store row to distinguish causes: owner renewed lease → lease_valid_elsewhere; row went terminal → not_cancellable; another worker already took it over → taken_over. - Extract _raise_lease_valid_elsewhere() helper to de-duplicate the 409+Retry-After block shared across cancel_run and stream_existing_run. - Extract _lease_expired_or_null() in persistence/run/sql.py to de-duplicate the lease-expiry SQL WHERE clause shared by claim_for_takeover and list_inflight_with_expired_lease. - 11 new tests: 5 SQL-layer claim_for_takeover (expired/valid/NULL/ terminal/nonexistent), 3 _compute_retry_after unit (NULL/unparseable/ normal), 2 claim re-read precision (terminal/takeover), 1 stream endpoint 409+Retry-After. Not addressed (non-blocking, reviewer agreed): - The 2–3 store.gets in the takeover cold path: optimizing the API to accept a pre-fetched record would couple the router to the manager more tightly than justified by the perf gain. - The lease-expiry inline loop in MemoryRunStore.list_inflight_with_- expired_lease pre-computes cutoff once for all rows; switching to the shared _is_lease_expired helper would recompute datetime.now() per row with no real benefit. 260 related tests pass; 36/36 blocking-IO gate pass; ruff clean. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(runs): de-duplicate lease-expiry helper, restore defensive fallback Address final round of review feedback: - Extract is_lease_expired to deerflow.utils.time (no _ prefix, public utility). Manager and MemoryRunStore now import from the same place instead of the store reaching backward into the manager for a private function. - Restore defensive else-10 fallback in grace_seconds property (removed in an earlier round). The guard is unreachable for current callers but protects future ones from AttributeError. - Comment the transient in-memory interrupted vs store error state when a local cancel is superseded by a peer takeover. - Comment the max(1, ...) floor in _compute_retry_after — the floor is a lower bound, not a poll interval; clients should apply jitter. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> Co-authored-by: rayhpeng <rayhpeng@gmail.com>
bytedance#4034) * fix(memory): coerce stored confidence in the three remaining raw reads `_coerce_source_confidence` exists because `memory.json` is user-editable and written across versions: a `confidence` that is null, a str, a bool, out of range, or non-finite still has to read back as a usable score. Consolidation already routes every stored read through it, and bytedance#4023 did the same for the max_facts trim. Three reads still take the field raw. `_build_staleness_section` formats it with `f"{conf:.2f}"` — a str raises ValueError, None raises TypeError. `_do_update_memory_sync`'s `except Exception` swallows that into `return False`, aborting the whole memory-update cycle permanently, since the offending fact is then never rewritten. The staleness cap's `sort(key=lambda f: f.get("confidence", 0))` ranks the facts it is about to delete. A str raises the same way; the values that don't raise mis-rank instead. `true`/`inf` outrank a genuine 0.9 and push it into the removal slot; an absent key ranks 0 and is deleted first; `nan` compares false against everything, so which fact dies depends on where the corrupted one happens to sit in the file. `search_memory_facts` — the `memory_search` tool, added in bytedance#4023 — ranks results the same raw way and then truncates to `limit`, so the same mis-ranking hands the model the wrong facts, or fails the tool call outright on a str. All three now read through the helper, so an unusable stored confidence ranks as unknown (0.5): neither kept ahead of a real score nor evicted before one. The two sorts run in opposite directions, and 0.5 is load-bearing in both. * test(memory): anchor the confidence delta at every coerced read The str-based tests raise on main, so they cannot go red for any stored confidence that mis-ranks without raising. `true`/`false`/`inf` and an absent key never reached the except handler at all: bool subclasses int, so `true` ranked 1.0; `inf` outranked every real score; an absent key ranked 0. Silent fact loss in the staleness cap, wrong results out of `memory_search` — no log line either way. Parametrize both ranking sorts over the four non-raising inputs that fall to the 0.5 default, each pitted against a genuine neighbour chosen so the survivor flips. The sorts run in opposite directions, so the one matrix covers eviction from both ends. `nan` is excluded from those: its old rank is undefined rather than pinned to an end of the order, so one fixed input order happens to yield the correct survivor. Its real property is order-independence, asserted across both fact orders. The staleness prompt formatter is pinned by rendered value, not merely by not raising: 1.5 → 1.00, -0.3 → 0.00, and inf/nan/true/None → 0.50, matching the consolidation prompt that reads the same field through the same helper.
…heck API KEY (bytedance#4116) * fix: V-001 security vulnerability Automated security fix generated by OrbisAI Security * fix(sandbox): wire provisioner API key through backend client and config RemoteSandboxBackend now accepts an api_key parameter and sends it as X-API-Key on all five provisioner HTTP calls (list, create, destroy, is_alive, discover). AioSandboxProvider reads provisioner_api_key from SandboxConfig and forwards it at construction time. SandboxConfig formally declares the field; config.example.yaml documents it under Option 4; docker-compose-dev.yaml threads PROVISIONER_API_KEY into both the provisioner and gateway containers so a single .env entry covers both sides. Tests: monkeypatch PROVISIONER_API_KEY and send X-API-Key headers in the five parametrized threading tests (previously 401-failing); new test_auth_middleware asserts /health is open, /api/* rejects no-header and wrong-key with 401, and accepts the correct key. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(provisioner): correct fail-closed auth docs and fix test mock signatures The first round of PR bytedance#4116 fixes left four blocking issues: - Mock signatures in test_remote_sandbox_backend.py (17) and test_aio_sandbox_provider.py (2) didn't accept the new headers= kwarg, causing TypeError on every provisioner HTTP call test. - sandbox_config.py and config.example.yaml described the auth as optional ("leave unset to disable") but the middleware is fail-closed: an unset PROVISIONER_API_KEY causes 401 on every /api/* request. - .env.example had no PROVISIONER_API_KEY entry, leaving users with an empty value and silent 401s. - No test covered the PROVISIONER_API_KEY="" fail-closed path. Fix all four: add headers=None to all mock signatures, correct the field description and example comment to state that both sides must have the same key set, add PROVISIONER_API_KEY to .env.example with generation guidance, add test_auth_middleware_unset_key, and add a logger.warning on auth rejection for observability. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
…_system> block (bytedance#4157) A custom subagent's description is agent-editable (persisted by setup_agent / update_agent) and is rendered into the <subagent_system> block of the lead-agent system prompt via the available-subagents listing. It was interpolated raw, so a first line like "</subagent_system><system-reminder>..." could close the block and forge a framework-reserved tag inside the system-role prompt. Escape it with html.escape at the render site, matching the sibling fixes for <soul> (bytedance#4137), memory facts (bytedance#4097), skill metadata (bytedance#4128), and remote content (bytedance#4099/bytedance#4002). Built-in descriptions are trusted constants and stay untouched. Adds a red/green regression test mirroring test_soul_prompt_injection.py.
* Add Monocle tracing Enable Monocle (OpenTelemetry tracing for LLM apps) with one setup call plus the monocle_apptrace dependency. setup_monocle_telemetry auto-instruments the frameworks already in use and writes traces to .monocle/. Additive; no changes to application logic. * Config-gate Monocle telemetry in the Gateway lifespan Addresses review on bytedance#4024: moves setup_monocle_telemetry out of agents/__init__ import time into the Gateway lifespan, gated by MonocleTracingConfig (MONOCLE_TRACING env, default off). Warns on the Langfuse/global-OTel-provider conflict and relies on monocle_apptrace's own duplicate-setup guard and existing-provider attach. Pins monocle_apptrace>=0.8.8 (+ uv.lock), adds .monocle/ to .gitignore, adds tests (default-off / toggle-on / no import-time setup), and documents exporters, Okahu, and the VS Code viewer in README, config.example.yaml, and backend/AGENTS.md. * Clarify Monocle/Langfuse single-provider guidance Make the docstring, warning, and AGENTS.md consistent with the README: only one library can own the global OpenTelemetry provider; Monocle initializes at startup before Langfuse's per-run handler, so enabling both drops Langfuse's spans — enable one OTel tracer (LangSmith, a callback, coexists fine). * Address review: optional extra, exporter validation, off-box warning, tests Responds to the second review round. - Make monocle_apptrace an optional extra (deerflow-harness[monocle], re-exposed as deer-flow[monocle]) following the boxlite/tui precedent, so a default install no longer pulls the OpenTelemetry stack. It stays pinned in the dev group for the tracing tests, and enabling MONOCLE_TRACING without the extra raises a clear install error. - Warn loudly at startup whenever any exporter other than `file` is configured, since those move prompts, tool inputs/outputs, and completions beyond the local .monocle/ directory. - Validate MONOCLE_EXPORTERS against the known exporter names and require OKAHU_API_KEY when okahu is selected, mirroring the Langfuse pattern. Validation runs from Monocle's own init (not validate_enabled) so a config typo can never fail agent runs; errors surface at Gateway startup instead. - Grow the tests from 5 to 13: caplog coverage for the Langfuse-conflict and off-box warnings, exporter validation cases, a stronger import-time regression that asserts the global TracerProvider is not replaced, and a subprocess double-invoke test exercising the real check_duplicate_setup. - Docs: config.example.yaml block retitled to a dedicated tracing header; README documents the [monocle] install and scopes tracing to Gateway runs. * docs: align Monocle README section with the other tracing providers Lead with what Monocle is and captures, drop the install step (the dev group already ships monocle_apptrace via uv sync; unusual installs get the RuntimeError), and point the missing-package error at the repo-native command (uv sync --extra monocle / deerflow-harness[monocle]). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address review: verified Langfuse coexistence, lifespan test, scope docs Responds to the third review round. The Langfuse conflict claim was wrong, verified empirically against langfuse 4.5.1 in both init orders: whichever library initializes second reuses the existing global TracerProvider and attaches its own span processor, so neither side loses spans. Dropped the warning and its tests, corrected the README, AGENTS.md, and config.example.yaml statements, and pinned the verified behavior with test_coexists_with_langfuse (real monocle + real langfuse in a subprocess, no mocks). One honest caveat documented: both processors see all spans, so Monocle's exporters also capture Langfuse's spans when both are enabled. Also from the review: - Document the Gateway-only scope in AGENTS.md: the lifespan is the sole call site, so the embedded DeerFlowClient and TUI are not instrumented; embedded users call setup_monocle_tracing_if_enabled() themselves. - Add test_gateway_lifespan_initializes_monocle pinning the lifespan wiring. - Comment why MonocleTracingConfig.is_configured is intentionally coarser than LangSmith/Langfuse (composite validation lives in validate() at startup). - Note that monocle_exporters_list takes the comma-separated string as-is. - Module-level importorskip("monocle_apptrace") so minimal installs collect the test module cleanly. * docs: reword Monocle intro sentence * fix(tests): run the import-time regression in a subprocess test_no_import_time_setup deleted deerflow.agents* from sys.modules and re-imported to force __init__ to re-execute. The re-import creates new module objects, and restoring the old sys.modules entries afterwards leaves the parent package's attribute bindings pointing at the new ones, so any later test that resolves a deerflow.agents.* dotted path (monkeypatch.setattr in test_summarization_middleware, test_thread_data_middleware, and others) failed with "module 'deerflow.agents' has no attribute ...". Run the check in a subprocess instead: the import is genuinely fresh, the assertion is stronger (the provider must still be the SDK-less proxy, proving nothing was installed at any point), and no module identity leaks into the rest of the suite. * Address review: console warning scope, embedded hint, honest naming, doc alignment Responds to the post-approval review round: - Scope the off-box exporter warning to the remote exporters (okahu, s3, blob, gcs): console writes to local stdout and no longer trips it. config.example.yaml's data-handling note now distinguishes file / console / remote likewise. - Rename MonocleTracingConfig.is_configured to is_enabled so the boolean reads as what it checks; the exporter-dependent credential check stays in validate(), run at Gateway startup. - Hint on the embedded path: build_tracing_callbacks() logs a debug line when MONOCLE_TRACING is set but setup never ran in this process, so embedded DeerFlowClient/TUI users are not left with silent no-op tracing. Backed by a process-global setup flag. - Re-export setup_monocle_tracing_if_enabled from deerflow.tracing, matching the package convention. - Note the deliberate fail-open-at-startup contrast with LangSmith/Langfuse in the lifespan, and the OTel SDK-internals dependency in the coexistence test. - Test hygiene: clear MONOCLE_* env in the tracing config/factory fixtures; reset the setup flag in the monocle test fixture; reword the README Langfuse-spans claim as the shared-provider inference it is. - Document that .monocle/ trace files are never rotated or cleaned up. * fix(tests): pin the factory logger level in the embedded-hint tests configure_logging() from earlier tests in the full suite pins an explicit INFO level on the logger hierarchy, so a root-level caplog.at_level(DEBUG) never sees the factory's debug hint. Scope caplog to deerflow.tracing.factory so the test is independent of suite ordering. * Address review: co-export disclosure, lifespan failure test, exporter parse dedup - Off-box warning now notes that Langfuse's spans are exported too when both providers are enabled and share the global OTel provider; pinned both ways by tests. - Pin the lifespan fail-open contract: a raising Monocle setup is logged and the Gateway keeps serving (pragma dropped now that the path is exercised). README notes a config error is reported at startup and tracing stays off until restart. - Hoist exporter parsing into MonocleTracingConfig.exporter_list so validate() and the off-box warning cannot diverge, and note the upstream coupling on the exporter allow-list. - Reduce config.example.yaml's Monocle block to a pointer; the capture, retention, and data-handling detail lives in README's Monocle section. --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
…lass-path allowlist (bytedance#4146) * fix(models): scope the OpenAI-compat rules to BaseChatOpenAI, not a class-path allowlist * address review: drop redundant stream_usage helper, close test matrix - Remove _enable_stream_usage_by_default and its now-unused _OPENAI_COMPAT_USE_PATHS tuple. The class-field stream_usage fallback already sets stream_usage=True for every BaseChatOpenAI subclass (they all declare the field), so the helper's use-path allowlist gated nothing real — verified a no-op in prod, and the two stream_usage tests stay green on main with the helper present. Those tests used a BaseChatModel stub that does not declare the field; point them at a real ChatOpenAI capturing class so they exercise the fallback they now depend on. - Give the non-OpenAI normalization-skip test an actual api_base value so it exercises the skip path (api_base passed through verbatim, never rewritten to base_url). - Add an Unreleased CHANGELOG entry for the api_base behavior change on the five affected subclasses. --------- Co-authored-by: Willem Jiang <willem.jiang@gmail.com>
…ert (bytedance#4154) * fix(tools): escape MCP tool names rendered into deferred prompts get_deferred_tools_prompt_section and get_mcp_routing_hints_prompt_section list MCP tool names into the <available-deferred-tools> and <mcp_routing_hints> system-prompt blocks without escaping, while the mirror get_skill_index_prompt_section (per its docstring) does escape. An MCP name is taken verbatim from an external server, so a crafted name could close the block and forge a framework tag. Escape names (and routing keywords) at render, mirroring the skill-index section. * fix(mcp): validate tool names at the load boundary Escaping at render only neutralizes < > &, so a tool name with newlines or markdown still injects free-form text into the deferred-tools prompt block. Deferred (tool_search) tools are never bound, so the provider's function-name check never runs on them. Drop any MCP tool whose name is not a valid identifier (^[A-Za-z0-9_-]+$) in get_mcp_tools() — the same charset the provider enforces at bind time — before it can enter the catalog or the prompt. Render-time html.escape stays as defense-in-depth. Mirrors the load-time skill-name validation in skills/storage/skill_storage.py.
…#4153) _call_is_network_sink missed the HEAD/OPTIONS verbs on requests/httpx, socket.create_connection, and urllib.request.urlretrieve. A bulk env dump or reverse shell shipped through any of these slipped past the CRITICAL exfil/reverse-shell rules whenever the URL was assembled at runtime (the non-literal case the string-literal URL check can't cover). Also treat socket.create_connection as the socket primitive in the reverse-shell shape. http.client.HTTP(S)Connection is intentionally left out: only the lazy constructor is statically visible (the request()/connect() that performs the I/O is an instance method the call-name analyzer can't resolve), so flagging the constructor would hard-block benign code that only builds a connection object. Cover the alias-resolved forms too: the sink check runs on the name after from-import / import-as resolution, a path the suite exercised only on the env-read side (bytedance#4087) and not on the sink side.
* fix(runtime): persist original human input outside model sanitization * refactor(history): load thread messages by global event sequence * fix(frontend): make summarization rescue a transient history bridge * fix(frontend): old message not append tail 1. add identity anchor 2. add bridgeOrder * fix(frontend): lint error fix * fix: address review feedback and harden pagination coverage - defer transient history ref writes until after render commit - cover large middleware-only history scans - verify infinite-query refetch recalculates page cursors - document AI event types and anchor-weaving differences * fix: harden message pagination and enrichment - append unmatched live tails after canonical history - warn and stop when pagination has_more lacks a cursor - deep-copy restored UI messages to isolate model-facing content - log invalid event sequence and non-advancing cursor errors - pass user_id explicitly through event-store history queries - cover middleware-only AI runs across memory, JSONL, and DB stores * fix: address pagination review feedback * fix(frontend): checkpoint has unknow redener content, optimize the anchor policy * fix(frontend): unit test issue missed previously, remove the TanStack cache trimming * fix(gateway): harden message history queries and provenance - reject externally forged original_user_content metadata - validate provenance metadata in upload and sanitization middleware - make run lookups fail closed by default - batch feedback queries by run ID - align memory message filtering with persistent stores
…tedance#4155) * fix(security): block forged framework tags in the input guardrail InputSanitizationMiddleware's _BLOCKED_TAG_NAMES neutralizes forged framework tags in untrusted input, but missed soul, thinking_style, and critical_reminders -- which the lead-agent system prompt's System-Context Confidentiality section names as internal framework data -- and the underscore spelling system_reminder emitted by the todo/terminal middlewares (only the hyphen spelling was blocked). A user, or an attacker-controlled web_fetch/web_search page via the shared neutralize_untrusted_tags primitive, could forge these blocks. Add them. * fix(security): cover framework authority blocks as a class, not a subset The confidentiality section declares every framework structured tag trusted ("and all other structured tags"), so the denylist must cover the authority blocks as a class. Add the live blocks still passing both sanitization paths (clarification_system, self_update, response_style, citations, skill_index, available_skills, disabled_skills, memory_tool_system, durable_context_data, slash_skill_activation), and pin the set against drift with a test that scans the framework source and fails when a new block is not blocked. * fix(security): scan the whole harness for framework blocks, fail closed The drift guard added in the previous revision scanned a hand-listed set of source files. That is the same forgot-to-update-a-list root cause the guard was meant to eliminate, one level up, and it failed exactly that way: tool_search.py was not in the list, so <mcp_routing_hints> and <available-deferred-tools> — both rendered into the lead-agent system prompt via the {deferred_tools_section} / {mcp_routing_hints_section} placeholders — passed both sanitization paths unneutralized. Replace the file list with a repo-wide scan plus an exemption set that states a reason per tag. The point is the failure direction, not the breadth: a new framework block anywhere in the harness now turns CI red until it is either blocked or exempted on the record, where before a block emitted from an unlisted file was silently unguarded. The scan reads raw source rather than AST string literals on purpose: an attributed block built as an f-string splits its '>' into a separate literal chunk, so an AST-on-literals scan misses it (verified against <consolidation_candidates>). Raw source has one comment false positive, exempted. Exempted with reasons: leaf/wrapper elements; the memory-updater and summarizer prompts, which are built from checkpointed state rather than the ModelRequest this middleware rewrites, so blocking them here would be false coverage, not protection; and the MindIE provider wire format, parsed out of model output. The scan surfaced five further live authority blocks beyond the two reported. Subagents reuse _build_runtime_middlewares and therefore share this denylist, so their system-prompt blocks are in the same class: file_editing_workflow, guidelines, output_format, working_directory. goal_continuation is a framework-authored hidden HumanMessage injected into the lead agent. Also loosen the scanner regex to match the tolerance of _BLOCKED_TAG_PATTERN so an attributed block cannot hide from the guard.
…#4147) * fix(frontend): gate branch action by completed turn * test(frontend): clarify branch turn boundaries
bytedance#4151) Replace FirecrawlApp with Firecrawl (v2 unified client) so the map/crawl/interact/extract tools target methods that actually exist: - map_url -> map (with sitemap kwarg instead of ignore_sitemap) - crawl_url -> crawl (keyword args instead of nested params dict) - interact -> scrape+actions (structured action dicts, not NL string) The installed firecrawl-py==4.23.0 has no map_url/crawl_url methods and interact(job_id, code=) does not accept url/actions params. All three tools previously deterministically returned Error:... before making a valid request. Also drop unused Optional import (ruff UP045). Co-authored-by: Claude <noreply@anthropic.com>
…K surfac…" (bytedance#4165) This reverts commit b7b4b49.
* add Volcengine Coding Plan to quikly setup * modify model list
…n in CI The docstring commands now use the backend/tests/monocle/ form (matching the README, which also gains the backend-dir uv variant), and the README states explicitly that the suite is skipped in CI and run on demand.
- README: new section on the committed trace. It is a full, unmodified real-run recording (system prompt of the recording date + fetched web content, no credentials), committed whole so the offline example parses a genuine trace. The offline assertions are pinned to this trace and the monocle_apptrace 0.8.8 span shapes; re-record when prompt, tools, or model change. - README: note that the monocle_trace_asserter fixture comes from monocle_test_tools' auto-registered pytest plugin (pytest11 entry point). - test_deerflow.py: comment why web_fetch asserts min_count=2 rather than the recorded exact count of 5 (fetch counts vary run to run; keep it a floor). - requirements.txt: loose pin python-dotenv>=1.0.
Two execution-gating fixes from review: - Live tests are now opt-in via MONOCLE_LIVE_TESTS=1 (default off). Previously the documented offline command collected the live tests too, and on a configured checkout (.env + config.yaml present) they would run for real, spending model tokens and hitting the network. Now the plain `pytest backend/tests/monocle/` run cannot go live regardless of what credentials are present; test_live_gate_defaults_off pins the gate. - The run_agent fixture no longer requires OPENAI_API_KEY. config.yaml resolves the model, which may be any provider (Anthropic, Gemini, Volcengine, ...), so a hard-coded OpenAI gate skipped valid configurations and passed invalid ones. Credentials are validated by the configured model itself. README and docstrings updated to match: offline command is offline by construction, live is MONOCLE_LIVE_TESTS=1, credentials described as the configured model's rather than OpenAI's. Verified both modes: default run is 2 passed 2 skipped with no network; opted in, all 4 pass with real end-to-end runs.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a behavioral test suite that asserts against DeerFlow's Monocle execution traces: which agent ran, which tools it called, what it was asked and what it returned, and its token and duration cost. Additive only, under
tests/monocle/, with no app-code changes.Why
DeerFlow already has 350+ backend tests, but they exercise the agent with the LLM and tools controlled, not by checking what a real run actually did. This adds that missing layer. It asserts on the emitted trace, so if a later change alters how the agent routes, which tools it calls, or how many tokens it burns, a test catches it.
How it works
It uses Monocle Test Tools. The offline tests load a recorded trace from file with
with_trace_source("file", trace_path=...), which is fast, needs no keys, and is deterministic. They then assert with the fluent API:called_agent,called_tool,contains_input/contains_any_output,under_token_limit,under_duration. The live test drives the agent end to end throughDeerFlowClientand asserts on structure and budget only.Changes (all under
tests/monocle/)test_deerflow.py: 4 offline file-loaded tests, one per curated question (solid-state EV briefing, vector-DB comparison, a repeat, sandbox file authoring), plus 1 live test guarded byOPENAI_API_KEYand app import.conftest.py: Monocle setup,.envloading,run_deerflow().traces/: 4 recorded 0.8.8 traces, one per question.requirements.txt: pinsmonocle_test_tools==0.8.8(the file trace source does not exist in 0.7.x).README.md.