This workspace uses a small execution-profile layer so shell actions are chosen deliberately instead of ad hoc.
-
inspect_local- Use for read-only inspection, diagnostics, and state review.
- Default for
rg,sed,git status, config reads, and log checks. - Privacy preflight is usually unnecessary unless the inspection output will be exported or handed off.
-
workspace_edit- Use for normal local edits inside the workspace.
- Good default for single-file skill updates, docs, and maintenance scripts.
-
risky_edit- Use when a change touches multiple files, automation wiring, routers, tasks, cron jobs, or reusable skills.
- A checkpoint is required before the command runs.
-
service_ops- Use for local scripts that also touch live services or APIs.
- Examples: calendar writes, Sheets writes, meeting-recording pipeline, outbound delivery hooks.
- Run privacy preflight when user/private context is passed to an external API.
-
remote_handoff- Use when the operation belongs on a remote host, SSH target, node-specific runtime, or container.
- Do not pretend this is a normal local shell step. State the target and handoff explicitly.
- Run privacy preflight before sending context to the remote target.
Pick the narrowest profile that matches the real risk:
- Read-only?
inspect_local - Local edit, low blast radius?
workspace_edit - Local edit, high blast radius?
risky_edit - Local script plus external side effects?
service_ops - Needs another machine/runtime?
remote_handoff
- Prefer
inspect_localbeforeworkspace_edit. - Prefer
workspace_editbeforerisky_edit. - Escalate to
risky_editwhen changing shared skills, scripts, routing layers, task automation, or environment-wide behavior. - If the right answer is
remote_handoff, say so early instead of silently faking local execution. - Name the real runtime target with
--runtime-targetwhenever the backend is not just the local workspace shell.
references/capability_boundaries.json maps the semantic lanes read_only, local_write, workflow_edit, external_send, and high_risk_control onto these existing profiles plus the action-scope gate. The mapping is intentionally not a second profile hierarchy. high_risk_control is disabled by default; inspect verbs remain locked to read_only regardless of older memory or recovered chat context.
Execution profiles are not prompt advice. They are the boundary where Helm turns model proposals into governed operations:
- the model proposes the next action
- the harness validates the schema, profile, command guard, tool grant, and skill contract
- the harness authorizes, blocks, or requests approval before execution
- the harness executes only the authorized path
- the harness records task state, guard decisions, checkpoint references, tool grants, validation evidence, and finalization state
- the harness returns observations, including denials, approval requirements, timeouts, errors, aborts, and handoff requirements
Every tool call or governed action must produce a result record. A denied, timed-out, failed, paused, or aborted action is still an observation; it should not disappear into a transcript-only explanation.
For long-running work, Helm also records resumable runtime state in
.helm/long-running-runtime.json:
- phase checkpoints preserve input hashes, processed items, pending items, output artifacts, tool evidence, and idempotency keys
- approval pauses preserve pending action context and a resume command
- specialist agents are declared through an agent registry with tool, memory, model, timeout, owner, version, and output-contract metadata
The task ledger remains the audit trail. The long-running runtime file is the control state used to resume from the last successful phase or continue after a human approval.
For workspace_edit and risky_edit, every changed line should trace directly to the user's request.
Agents should not:
- refactor adjacent code unless explicitly requested
- rewrite comments or formatting unrelated to the task
- add speculative abstractions or configurability
- remove pre-existing dead code unless asked
Agents may remove only the unused imports, variables, or helpers introduced by their own change.
Helm's current edit policy implements a patch-first helper:
references/edit_policy.jsonsetsdefaulttopatch_firstscripts/edit_policy.pytracks per-file patch failures and recommendsreload_context_then_decomposeafter repeated failure- the policy can require checkpoints for target kinds such as
shared_workflow,skill_router, andautomation
SmallCode-style read-before-write now has a deterministic policy surface in
Helm. scripts/edit_policy.py validates read evidence before mutation, treats
missing or stale path/mtime/size evidence as a blocker, and keeps whole-file
rewrites limited to new files, generated artifacts, small files, or explicit
user requests.
Execution is not the whole task boundary.
After the command or handoff path ends, Helm should still decide whether the result needs durable state capture.
A task is not complete merely because files changed. It is complete only when the intended outcome has a named verification gate. Examples:
- bug fix: regression test reproduces the bug and passes
- documentation change: links, anchors, or direct inspection validate the rendered guidance
- package or release change: metadata parses and install/build checks pass
- workflow change: dry-run, static validation, or postflight evidence passes
- Obsidian artifact change: Markdown, Base, or Canvas structure is checked according to the artifact type
Examples:
- repo docs, workflow rules, release actions, or reusable scripts changed
- live service or integration behavior changed
- note, memory, ontology, or other durable knowledge sources changed
The profiled runner now writes a memory_capture plan into the final task-ledger state so this decision is visible instead of implicit.
Completion claims require evidence. A final answer, assistant message, or
compacted summary is not sufficient evidence by itself. The durable record
should point to task evidence such as exit code, diff inspection, test/lint
output, provider result, checkpoint id, write validation, cleanup evidence, or
explicit completion_evidence.
For conversation-only or synthetic task paths, the same rule still applies: auditability should use an explicit lifecycle instead of a single terminal row.
The preferred ledger shape is:
queuedrunning- final state such as
completed,failed, orhandoff_required
That lifecycle keeps timestamp audits, failure review, and rollback reasoning aligned across shell-backed and conversation-backed execution.
Execution profiles should make private-data boundary decisions explicit.
Use helm privacy scan for a no-write check and helm privacy tokenize when private text must cross a boundary in recoverable form. The default vault and audit log live under the workspace state directory.
Recommended defaults:
inspect_local: no preflight unless output will be exported or sharedworkspace_edit: scan before writing user/private context into durable docs or fixturesrisky_edit: scan checkpoint/state material when it may contain raw private contextservice_ops: tokenize user/private context before external API calls when the raw value is not requiredremote_handoff: tokenize context before handoff; restore only on the authorized local boundary
Secrets such as API keys, passwords, access tokens, and refresh tokens should be redacted instead of stored as recoverable vault entries.
See Privacy Boundary.
-
List or inspect profiles:
python3 ~/Helm/scripts/run_with_profile.py listpython3 ~/Helm/scripts/run_with_profile.py show risky_editpython3 ~/Helm/scripts/run_with_profile.py policypython3 ~/Helm/scripts/run_with_profile.py validate-manifests --jsonpython3 ~/Helm/scripts/run_with_profile.py audit-manifest-quality --json
-
Run a command with a declared profile:
python3 ~/Helm/scripts/run_with_profile.py run workspace_edit -- git -C ~/Helm status --shortpython3 ~/Helm/scripts/run_with_profile.py run service_ops --task-name "meeting pipeline" -- python3 /path/to/helper.pypython3 ~/Helm/scripts/run_with_profile.py run remote_handoff --runtime-target ssh:gpu-box --runtime-note "Docker build belongs on remote builder" -- docker build .
-
Create a checkpoint directly:
python3 ~/Helm/scripts/workspace_checkpoint.py create --label risky-router-edit --path examples/demo-workspace/skill_drafts/router-context-demo --path scriptspython3 ~/Helm/scripts/workspace_checkpoint.py preview <checkpoint-id>
-
Inspect the task ledger:
python3 ~/Helm/scripts/run_with_profile.py ledger --limit 20python3 ~/Helm/scripts/run_with_profile.py rollback --task-id <task-id> --jsonpython3 ~/Helm/scripts/task_ledger_report.py --summarypython3 ~/Helm/scripts/task_ledger_report.py --failed-only --limit 20python3 ~/Helm/scripts/task_ledger_report.py --skill router-context-demo --summarypython3 ~/Helm/scripts/task_ledger_report.py --latest --summary
-
Inspect low-level command execution:
python3 ~/Helm/scripts/command_log_report.py --summarypython3 ~/Helm/scripts/command_log_report.py --component router-context-demo --failed-only
risky_editautomatically creates a checkpoint before execution.risky_editstores the createdcheckpoint_idin later task-ledger states when checkpoint creation succeeds.- guard
require_approvaldecisions create a runtime approval pause before the runner exits withEXIT_GUARD_REQUIRE_APPROVAL. remote_handoffrecords a handoff task instead of pretending to execute locally, and requires--runtime-target.- If
--skillis provided, the runner checks the skill-localcontract.jsonmanifest first and rejects disallowed profile/skill combinations. service_opsruns are appended to.helm/task-ledger.jsonlso detached or side-effectful work is auditable later.- Final task-ledger states include a visible
memory_captureassessment so operational completion is inspectable. - Intentional weak-model or small-model fallback paths should be documented as explicit operating exceptions, not silently normalized away in the runner.