This phase upgrades OpenBaD from a biologically inspired agent framework into a persistent autonomous execution system built on top of the existing substrate.
OpenBaD already has strong foundational systems:
- An MQTT-backed nervous system
- A reflex arc and finite state machine
- Active inference and surprise-driven scanning
- Endocrine regulation
- Cognitive routing and context budgeting
- Stratified memory and sleep consolidation
- A web UI and operator-facing control surface
What OpenBaD does not yet have is a complete, restart-safe autonomous task execution layer with persistent scheduling, explicit research escalation, isolated external tool use, and a formal capability registry.
This phase defines how to add those systems without replacing the architecture that already exists.
Phase 9 must deliver the following capabilities:
- A persistent task system backed by SQLite.
- A heartbeat-driven scheduler that emits semantic work signals instead of embedding planning in timer loops.
- A DAG-based task execution model with node dependencies, retries, blocking, and resumability.
- A research escalation stack for blocked or uncertain work.
- A trusted in-process capability registry for core OpenBaD actions.
- An isolated MCP bridge for third-party tools with no ambient access from System 1 paths.
- L2HR-backed reward evaluation at the task-node level.
- Full operator visibility through the WUI, telemetry, and auditable persistence.
The first implementation of this phase must not attempt the following:
- A full HDDL parser or external planner runtime.
- Distributed task execution across multiple nodes.
- General multi-agent collaboration.
- Arbitrary unrestricted code generation for reward programs.
- Replacement of the current memory subsystem.
- Replacement of the current active inference engine.
- Replacement of the current endocrine controller.
This phase extends the current implementation and must not re-architect working systems without a direct need.
Use the existing active inference infrastructure as the heartbeat-adjacent observation layer:
src/openbad/active_inference/engine.pysrc/openbad/active_inference/background_scanner.pysrc/openbad/active_inference/budget.pysrc/openbad/active_inference/insight_queue.pysrc/openbad/active_inference/plugin_loader.py
Reuse the current hormone controller and extend it with task-aware triggers:
src/openbad/endocrine/controller.pysrc/openbad/endocrine/l2hr.pysrc/openbad/endocrine/telemetry.py
Reuse routing, fallback, provider health, and context budgets:
src/openbad/cognitive/model_router.pysrc/openbad/cognitive/context_manager.pysrc/openbad/cognitive/config.py
Reuse existing topics, typed publish-subscribe patterns, and state transitions:
src/openbad/nervous_system/client.pysrc/openbad/nervous_system/topics.pysrc/openbad/reflex_arc/fsm.pysrc/openbad/daemon.py
Reuse the current memory controller and sleep pipeline:
src/openbad/memory/controller.pysrc/openbad/memory/base.pysrc/openbad/memory/episodic.pysrc/openbad/memory/semantic.pysrc/openbad/memory/procedural.pysrc/openbad/memory/sleep/
Extend the current server and front-end instead of creating a second operator interface:
src/openbad/wui/server.pysrc/openbad/wui/bridge.pywui-svelte/src/routes/
The implementation for this phase must obey these constraints.
System 1 includes the heartbeat path, reflex handlers, and low-cost monitoring loops. These paths must remain local-first, highly constrained, and denied external tool access.
High-cost reasoning, external tool execution, and research sessions must only happen in explicit task or research contexts.
The heartbeat, background scanner, and reflex arc must not inherit MCP tool access implicitly.
Task state, leases, heartbeat metadata, research queue entries, and audit logs must survive process restarts.
Every autonomous action must be attributable to a task, task node, research node, or operator command.
Rate limits, token budgets, task concurrency, and emergency suppression must be enforced by the runtime and not delegated to model compliance.
This phase is complete only when all of the following exist.
- SQLite-backed task and research state.
- Task CRUD and task execution APIs.
- A scheduler that emits work events based on persisted state.
- DAG execution with node lifecycle tracking.
- A trusted capability manifest and registry.
- An MCP bridge with task-scoped sessions and audit logs.
- Reward programs with task-node evaluation.
- Research queue prioritization for blocked work.
- WUI visibility into tasks, nodes, research, rewards, and scheduler state.
- Tests that validate restart safety, lease contention, isolation, and task transitions.
Phase 9 introduces four new first-class subsystems:
- The task subsystem
- The scheduler subsystem
- The capability subsystem
- The MCP subsystem
These subsystems must integrate cleanly with the existing cognitive, endocrine, active inference, and memory systems.
Create a new package:
src/openbad/tasks/__init__.pysrc/openbad/tasks/models.pysrc/openbad/tasks/store.pysrc/openbad/tasks/planner.pysrc/openbad/tasks/executor.pysrc/openbad/tasks/scheduler.pysrc/openbad/tasks/rewards.pysrc/openbad/tasks/research.pysrc/openbad/tasks/service.py
The task subsystem owns:
- Task creation and persistence
- Task decomposition into nodes
- Node dependency graphs
- Task leasing and concurrency
- Node execution and retry policy
- Task event logging
- Task note compaction
- Research escalation
- Reward program binding and evaluation
Define the following enums in models.py.
pendingreadyrunningblockedwaitingcompletedfailedcancelled
immediateshortmediumlong
user_requestedrecurringheartbeat_spawnedresearch_spawnedreflex_spawned
reasoncapabilitymcp_toolsummarizeverifywaitresearchconsolidate
The initial design must be strongly typed.
@dataclass(slots=True)
class Task:
task_id: str
title: str
description: str
kind: TaskKind
horizon: TaskHorizon
priority: int
status: TaskStatus
created_at: datetime
updated_at: datetime
due_at: datetime | None
parent_task_id: str | None
root_task_id: str
owner: str
lease_owner: str | None
recurrence_rule: str | None
requires_context: bool
isolated_execution: bool
notes_path: str | None@dataclass(slots=True)
class TaskNode:
node_id: str
task_id: str
title: str
node_type: TaskNodeType
status: TaskStatus
depends_on: tuple[str, ...]
capability_requirements: tuple[str, ...]
model_requirements: tuple[str, ...]
reward_program_id: str | None
expected_info_gain: float
blockage_score: float
retry_count: int
max_retries: int@dataclass(slots=True)
class TaskRun:
run_id: str
task_id: str
node_id: str | None
started_at: datetime
finished_at: datetime | None
status: TaskStatus
actor: str
routing_provider: str | None
routing_model: str | None@dataclass(slots=True)
class TaskEvent:
event_id: str
task_id: str
node_id: str | None
event_type: str
created_at: datetime
payload: dict[str, Any]The internal representation must be DAG-native, even if future HDDL support is added later.
The planner must emit:
- Nodes
- Dependency edges
- Execution hints
- Capability requirements
- Routing hints
- Retry policy
- Reward bindings
Do not block this phase on formal HDDL parsing.
Add a new package:
src/openbad/state/__init__.pysrc/openbad/state/db.pysrc/openbad/state/migrations/
Use a local SQLite database at:
data/state.db
The first migration must create at least the following tables.
taskstask_nodestask_edgestask_runstask_eventstask_notestask_leasesheartbeat_stateresearch_nodesresearch_findingsreward_programsmcp_auditscheduler_windows
The store must support:
- Atomic lease acquisition
- Recovery after process restart
- Efficient polling for due or ready work
- Append-only task event logs
- Idempotent scheduler wake handling
- Event ordering by durable timestamps
Leases are required to prevent duplicate task execution.
Each lease record must include:
lease_idowner_idresource_typeresource_idleased_atexpires_at
The executor must renew leases while work is ongoing. Expired leases must be reclaimable.
The scheduler must not embed task planning logic in a timer loop. It should emit semantic work signals based on persisted state.
src/openbad/tasks/scheduler.pysrc/openbad/tasks/heartbeat.py
The scheduler must:
- Wake at configured intervals
- Read
heartbeat_statefrom SQLite - Determine whether recurring work, blocked review, research review, or maintenance work is due
- Publish work events to the nervous system or invoke the task service directly
- Avoid duplicate wakeups for already leased work
- Respect cortisol, adrenaline, quiet hours, and sleep windows
OpenBaD currently uses the agent/... namespace in src/openbad/nervous_system/topics.py. Phase 9 must stay within that namespace.
Add the following topics:
agent/tasks/context_requiredagent/tasks/isolatedagent/tasks/eventsagent/research/deep_diveagent/scheduler/wakeagent/scheduler/sleep_windowagent/scheduler/maintenance
If topic templates are needed, define them in topics.py with the same pattern as existing topic constants.
The heartbeat loop must follow this algorithm.
- Wake on interval.
- Read persisted scheduler state.
- Query for due recurring tasks.
- Query for blocked tasks eligible for re-evaluation.
- Query for research nodes awaiting work.
- Query for maintenance or consolidation windows.
- If no work is due, increment a silent skip counter and return.
- If work is due, publish the appropriate event or dispatch to the task service.
The heartbeat_state table must track at least:
last_heartbeat_atlast_triage_atlast_context_required_dispatch_atlast_research_review_atlast_sleep_cycle_atlast_maintenance_atsilent_skip_count
The current observation plugin model is not sufficient for trusted action execution. Add a separate capability system.
src/openbad/capabilities/__init__.pysrc/openbad/capabilities/base.pysrc/openbad/capabilities/manifest.pysrc/openbad/capabilities/registry.pysrc/openbad/capabilities/executor.pysrc/openbad/capabilities/core_triage.py
Trusted in-process capability plugins must be described by openbad.plugin.json files.
Example manifest:
{
"id": "openbad-core-triage",
"name": "OpenBaD Core Triage",
"version": "1.0.0",
"tier": "trusted",
"kind": "tool",
"module": "openbad.capabilities.core_triage",
"capabilities": [
"create_task",
"queue_research",
"pause_task",
"resume_task",
"mark_task_blocked"
],
"permissions": [
"db.insert",
"db.update",
"mqtt.publish"
]
}The loader must enforce these rules.
- Only approved local directories may be scanned.
- Only package-local Python modules may be imported as trusted.
- Import must be side-effect free.
- Invalid manifests must fail closed.
- Permissions must be validated against
config/permissions.yaml.
class Capability(Protocol):
capability_id: str
async def execute(
self,
request: CapabilityRequest,
context: CapabilityContext,
) -> CapabilityResult:
...Required types:
CapabilityRequestCapabilityContextCapabilityResultCapabilityDescriptorCapabilityError
CapabilityContext must include at least:
task_idrun_idactorpermission_scopebudget_snapshotendocrine_snapshotmemory_controllermqtt_clientcancellation_token
The first implementation must provide the following trusted capabilities.
create_taskqueue_researchpause_taskresume_taskcancel_taskappend_task_notepublish_eventrequest_escalation
These capabilities are enough to let System 1 trigger or shape work without granting third-party tool use.
Third-party tools must be isolated from heartbeat and reflex paths.
src/openbad/mcp/__init__.pysrc/openbad/mcp/bridge.pysrc/openbad/mcp/session.pysrc/openbad/mcp/policy.py
These rules are mandatory.
- The heartbeat path has no MCP access.
- The background scanner has no MCP access.
- Reflex handlers have no MCP access.
- Task executor may create MCP sessions for nodes that explicitly require them.
- Research executor may create MCP sessions only if policy permits.
@dataclass(frozen=True)
class MCPPolicy:
allowed_servers: tuple[str, ...]
allowed_tools: tuple[str, ...]
max_calls: int
max_duration_seconds: int
allow_network_egress: boolEach MCP session must be:
- Created explicitly by the task or research executor
- Bound to one task run or research run
- Audited per tool call
- Torn down at the end of the run
- Denied access to undeclared tools
Every MCP invocation must create an audit record containing:
audit_idtask_idrun_idtool_nameserver_namestarted_atfinished_atstatusinput_summaryoutput_summaryerror_summary
The planner converts a task request into a DAG of executable nodes.
src/openbad/tasks/planner.py
The first version should be deterministic and template-driven with optional LLM refinement.
The planner must:
- Parse task intent
- Determine horizon
- Estimate node sequence
- Attach dependencies
- Mark context-required vs isolated nodes
- Estimate capability and model requirements
- Attach reward program templates
- Set retries and blocking thresholds
The planner must emit a structure that includes:
- Task metadata
- Nodes
- Edges
- Execution hints
- Default reward templates
- Initial research eligibility
The executor runs ready nodes under leases and updates state durably.
src/openbad/tasks/executor.py
The executor must:
- Lease a ready node
- Create a task run record
- Build a bounded execution context
- Route reasoning through
ModelRouterwhen necessary - Invoke trusted capabilities or MCP sessions only if declared
- Write structured notes and event records
- Evaluate reward programs
- Update endocrine hooks when appropriate
- Transition node and task states
The executor must not keep raw tool output in active prompt state beyond its immediate utility.
For medium and long horizon tasks:
- Raw output may be used during the current node.
- After node completion, replace raw output in working state with:
- A summary
- Extracted facts
- Follow-up implications
- Artifact references
- Persist detailed artifacts to disk when configured.
This must be implemented as a task-aware extension of the existing context budget logic in src/openbad/cognitive/context_manager.py.
On node failure, the executor must:
- Increment retry count
- Capture structured failure summary
- Update blockage score
- Either retry, block, or escalate to research based on policy
Research is a specialized escalation path for uncertainty and blockage.
src/openbad/tasks/research.py
@dataclass(slots=True)
class ResearchNode:
research_id: str
source_task_id: str
source_node_id: str
trigger_reason: str
blockage_score: float
expected_info_gain: float
urgency_score: float
priority_score: float
status: TaskStatus
findings_summary: str | None
artifact_path: str | NoneResearch queue priority should be computed as:
Default weights:
$w_b = 0.4$ $w_i = 0.3$ $w_u = 0.2$ $w_c = 0.1$
- A node becomes blocked or enters uncertain completion.
- The executor computes blockage and expected information gain.
- If thresholds are exceeded, a research node is queued.
- The research scheduler acquires the highest-priority item.
- The research run may use MCP if policy allows.
- Findings are summarized and persisted.
- Findings are written into episodic or semantic memory.
- The source task is re-evaluated.
Research findings must integrate with the existing memory system.
- Episodic memory stores trace-like records.
- Semantic memory stores reusable findings.
- Procedural memory is updated only when a repeatable workflow is verified.
OpenBaD already contains a natural-language-to-hormone mapper in src/openbad/endocrine/l2hr.py. Phase 9 must extend this into true task-node reward evaluation.
src/openbad/tasks/rewards.py
class RewardProgram(Protocol):
def evaluate(self, trace: ExecutionTrace) -> RewardResult:
...Required types:
ExecutionTraceRewardResult
ExecutionTrace must include:
task_idnode_idduration_msretriesapi_callsmcp_callstokens_usedbudget_remainingblockedcompletedverification_passedoperator_interruptendocrine_snapshot
RewardResult must include:
scalar_rewardhormone_adjustmentreasons
def evaluate(trace: ExecutionTrace) -> RewardResult:
reward = 0
reasons: list[str] = []
if trace.completed:
reward += 10
reasons.append("completed")
if trace.verification_passed:
reward += 5
reasons.append("verification_passed")
if trace.api_calls > 5:
reward -= 100
reasons.append("api_limit_exceeded")
if trace.blocked:
reward -= 15
reasons.append("blocked")
hormone_adjustment = {
"dopamine": 0.1 if reward > 0 else 0.0,
"cortisol": 0.1 if reward < 0 else 0.0,
}
return RewardResult(
scalar_reward=reward,
hormone_adjustment=hormone_adjustment,
reasons=reasons,
)Generated reward programs must not be arbitrary unrestricted code.
Allowed first implementations:
- Restricted Python subset
- Deterministic rule templates
- Validated declarative rules compiled to Python
This phase must connect task execution and research pressure to the endocrine system.
Increase cortisol when any of the following occur:
- Token budget exhaustion
- Thermal threshold breach
- Provider failure storm
- MCP rate-limit exhaustion
- Repeated node retries
- Excess blocked tasks
Effects:
- Suppress research branching
- Lower task concurrency
- Prefer cheaper routes in
ModelRouter - Defer maintenance and non-urgent work
Increase adrenaline when any of the following occur:
- Critical immune alert
- Explicit urgent user request
- Near-term deadline breach risk
- Cascading critical task failure
Effects:
- Suspend background research
- Widen context allowance for critical tasks
- Allow temporary soft-cap override
- Enter emergency scheduling mode
Increase dopamine when:
- A task completes with verification
- A high-value research node resolves useful uncertainty
- A repeatable workflow is verified for procedural storage
Increase endorphin when:
- Stress resolves after adrenaline or cortisol spikes
- Maintenance completes cleanly
- Consolidation or sleep windows complete successfully
Add a new file:
config/tasks.yaml
Initial schema:
tasks:
enabled: true
db_path: data/state.db
heartbeat_interval_seconds: 180
recurring_scan_interval_seconds: 300
blocked_review_interval_seconds: 600
research_review_interval_seconds: 900
max_concurrent_runs: 2
default_node_max_retries: 2
blocked_threshold: 0.65
research_threshold: 0.70
quiet_hours_start: "23:00"
quiet_hours_end: "06:00"
maintenance_window_start: "02:00"
maintenance_window_duration_minutes: 90
compaction:
medium_horizon_drop_raw_tool_output: true
long_horizon_drop_raw_tool_output: true
store_artifacts_on_disk: true
mcp:
enabled: true
default_max_calls: 5
default_max_duration_seconds: 300Optionally extend config/endocrine.yaml with task-related increments:
task_retry_cortisol_incrementresearch_success_dopamine_incrementdeadline_adrenaline_increment
The task system must be operator-visible from the start.
Add endpoints to src/openbad/wui/server.py.
GET /api/tasksPOST /api/tasksGET /api/tasks/{task_id}POST /api/tasks/{task_id}/pausePOST /api/tasks/{task_id}/resumePOST /api/tasks/{task_id}/cancelGET /api/tasks/{task_id}/eventsGET /api/researchGET /api/research/{research_id}GET /api/capabilitiesGET /api/mcp/auditGET /api/scheduler/state
Add WUI views for:
- Task list
- Task detail and node graph
- Research queue
- Capability inventory
- Scheduler state
- Reward evaluation traces
- MCP audit records
The UI does not need to be visually complete in the first pass, but it must expose enough state for debugging and operator trust.
If cross-process transport is required for task and research events, add new protobuf schemas under src/openbad/nervous_system/schemas/.
Potential files:
task.protoresearch.protoscheduler.protocapability.proto
Potential messages:
TaskCreatedTaskUpdatedTaskNodeUpdatedTaskRunStartedTaskRunFinishedResearchQueuedResearchResolvedSchedulerWakeCapabilityExecutedMCPAuditRecord
Do not add protobuf messages unnecessarily if SQLite plus HTTP is sufficient for the first implementation.
Implement Phase 9 in this order.
Build SQLite state, migrations, task models, and task creation APIs.
Acceptance criteria:
- Tasks persist across restart.
- Tasks can be created and queried.
- Task events are stored durably.
Build the heartbeat scheduler and lease acquisition.
Acceptance criteria:
- Due tasks are dispatched after restart.
- Duplicate dispatch is prevented by leases.
- Quiet hours and maintenance windows are respected.
Add the planner and executor with node transitions.
Acceptance criteria:
- Dependent nodes run in order.
- Downstream nodes remain blocked if upstream nodes fail.
- Retry and blocking behavior are deterministic.
Add the capability manifest, registry, and core triage capability pack.
Acceptance criteria:
- Manifests validate correctly.
- Capability inventory is exposed to operators.
- System 1 only sees the restricted capability set.
Add the MCP bridge and task-scoped sessions.
Acceptance criteria:
- Heartbeat and reflex paths have no MCP access.
- MCP sessions are task-scoped and audited.
- Tool access is denied unless declared by policy.
Add blocked-task research escalation and reward evaluation.
Acceptance criteria:
- Blocked nodes can enqueue research work.
- Reward traces are stored and inspectable.
- Research findings can feed back into task execution.
Connect task signals into endocrine behavior and expose everything through the WUI.
Acceptance criteria:
- Cortisol suppresses exploratory work under stress.
- Adrenaline prioritizes urgent work.
- Operator can inspect tasks, research, reward traces, and scheduler state.
Add at least the following tests.
tests/test_task_store.pytests/test_task_scheduler.pytests/test_task_executor.pytests/test_task_planner.pytests/test_capability_registry.pytests/test_mcp_bridge.pytests/test_reward_programs.pytests/test_research_queue.pytests/test_task_api.py
Minimum behavioral coverage:
- SQLite migration correctness
- Lease contention and expiry
- Duplicate heartbeat suppression
- DAG ordering and dependency blocking
- Retry and escalation thresholds
- Compaction of raw tool outputs
- Manifest validation and permission enforcement
- MCP isolation and audit logging
- Reward evaluation correctness
- Endocrine adjustments from task traces
An agent implementing this phase should follow these rules.
- Reuse OpenBaD's current nervous system, cognitive router, endocrine controller, memory controller, and WUI.
- Do not redesign working modules to fit a cleaner abstract architecture.
- Add persistence and orchestration first, then add richer autonomy.
- Keep System 1 narrow and non-ambient.
- Treat MCP as a scoped execution privilege, not a globally visible tool inventory.
- Prefer deterministic first implementations over speculative generality.
- Make every autonomous step replayable and inspectable.
Phase 9 is done when OpenBaD can:
- Persist tasks and task DAGs in SQLite.
- Resume scheduled and in-progress work after restart.
- Execute task nodes under leases without duplicate processing.
- Escalate blocked work into a research queue.
- Use trusted core capabilities without exposing external tools to heartbeat paths.
- Use MCP tools only in explicitly scoped task or research sessions.
- Evaluate task-node reward programs and translate outcomes into endocrine adjustments.
- Expose task, research, scheduler, capability, and audit state through the WUI.
- Pass the Phase 9 test suite.
At that point, OpenBaD will have crossed from biologically inspired reactive architecture into persistent, inspectable, and resource-governed autonomous execution.