Klebb ships with a chat widget that can write to cards and answer questions about your data. The widget is agent-agnostic: it talks to an an OpenAI-compatible chat-completions endpoint, which can be backed by whatever model you configure.
This doc covers:
- How the chat widget is configured
- How server-to-server writes from an external agent work
- What's NOT included (and what you'd need to build)
The in-page chat widget (health-chat.js) posts user messages to
whatever endpoint you configure and renders the response. Klebb speaks
the OpenAI chat-completions shape, so any endpoint that accepts that
shape works: a self-hosted gateway (LiteLLM, or similar), a cloud
provider's OpenAI-compat endpoint (your provider, Groq, Together,
DeepInfra), a local runtime (Ollama, vLLM, llama.cpp), or OpenAI /
OpenRouter directly.
Configure via environment variables:
| Var | Default | Purpose |
|---|---|---|
CHAT_ENDPOINT_URL |
— | Full URL of the chat-completions endpoint, e.g. https://api.openai.com/v1/chat/completions |
CHAT_API_KEY |
— | Bearer token sent as Authorization: Bearer <key> |
CHAT_MODEL |
— | Model name the endpoint expects (e.g. gpt-4o-mini, llama3.1, whatever your provider returns from its models list) |
HEALTH_SYSTEM_PROMPT |
built-in | System prompt sent with each turn |
CHAT_AGENT_NAME |
Chat |
Display name shown in the chat UI |
CHAT_AGENT_EMOJI |
💬 |
Emoji/char shown as the agent avatar |
The URL scheme (http:// vs https://) picks the transport. Host, port,
and path all come from the URL, so the endpoint doesn't have to live at
/v1/chat/completions — point at whatever path your provider uses.
If CHAT_ENDPOINT_URL is unset, the chat widget is disabled and
POST /api/chat returns 503.
The server passes a small set of tools to the chat-completions
endpoint on every turn. If your model supports tool-use, the loop in
chat/tools.js dispatches each call inline; if your model doesn't,
the tools are simply unused.
| Tool | Purpose |
|---|---|
create_manifest / delete_manifest / patch_manifest |
Author and edit cards |
read_manifest / list_manifests / write_manifest_data |
Inspect and update card data |
read_manifest_meta / read_manifest_rows / append_row / update_row / remove_row / reorder_rows |
Targeted reads + row-level mutations |
hide_card / show_card |
Master enable/disable |
set_notification / remove_notification |
Add, update, or remove a Web Push reminder on a card. set_notification is idempotent by (card_id, notification_id); remove_notification requires one-shot user confirmation. v1 trigger types: daily and weekly. The validator enforces title <= 30, body <= 80, label <= 80, items[] <= 10 per card; the system prompt forbids including numerical values or past-entry content in the body (notifications are reminders to act, not summaries). |
read_doc |
Fetch any allowlisted in-repo doc (README, MANIFEST-SCHEMA, this file, etc.) |
read_report |
Fetch any ingested report from $HEALTH_HOME/reports/. The agent gets the catalogue automatically in its system prompt; see REPORTS.md for how reports get there. |
get_recent_activity |
One-pass recency summary of every card (rowCount, lastEntryDate, ageDays, lastNDelta). The agent calls it before answering "how's my tracking" questions and before authoring a card (to match sibling conventions). |
hygiene_scan |
On-demand dashboard health check: stale / oversized / orphaned-input findings. Report-only; never mutates. stale is opt-in per card: only cards declaring meta.cadence.expectDays are ever reported, so a card with no declared cadence being quiet is not a finding. |
validate_manifest |
Dry-run a candidate manifest (no write). Returns {ok} or {ok:false, errors:[{path,message}]}. The system prompt directs the agent to call it before every create/patch. |
note_feedback |
Logs an anonymised bug report (kind: 'bug') or unmet-capability request (kind: 'feature') to data/_meta/feedback.jsonl. Paraphrased intent only, never user data. The in-app feedback form (drawer footer) writes the same log via POST /api/feedback, and the operator collects it via GET /api/admin/feedback (admin bearer, ?since=<ISO> cursor). |
The system prompt explicitly tells the model to refuse fast when the
user's request can't be carried out by any of the available tools in a
single generation. This is the answer to "the model tried to fudge a
reorder through write_manifest_data, the gateway timed out at 180s,
the user saw three minutes of dead air".
The standard refusal copy is one or two short sentences:
I can't do that in one step right now: <one-line reason>. <Optional: name the closest workaround the user CAN do, or what tool would be needed.>
Examples:
- "I can't reorder rows in one step right now: there's no reorder primitive, and the only tool that could do it would have to rewrite the whole data block (which times out on cards this size). You can re-order this card by editing the manifest file directly for now."
- "I can't merge two cards in one call: there's no cross-card transaction tool. I can copy rows from one to the other if you read them out yourself first."
Any future tool addition should keep the refusal pattern and tighten
it: when a new primitive lands (e.g. reorder_rows), the refusal text
for that intent stops applying and the agent should reach for the new
tool instead.
The agent loop has three budgets, all env-tunable:
CHAT_MAX_TURNS(default36): gateway round-trips per turn. One round-trip may batch several tool calls, but the prompt's own validate-before-create / read-before-append workflow means multi-card requests legitimately need many round-trips. When the cap is hit the reply keeps any progress text the model produced, appends how to resume ("keep going" works because the client resends the transcript), and carriescapped: truefor the client.CHAT_ITER_TIMEOUT_MS(default180000,0disables): soft per-iteration budget under the transport's hard 540s per-hop ceiling. A single step running past it aborts the in-flight gateway call and answers with timeout copy (HTTP 200), emitting[chat:<id>] iter=N gw=<ms>ms iter_timeoutin debug logs. Keep it strictly below the transport ceiling or it stops being soft and the turn errors 504 instead.CHAT_TURN_DEADLINE_MS(default720000,0disables): total wall clock for the whole turn. Without it, a raised iteration cap could stack per-step timeouts into a multi-minute silent spinner. The loop stops starting new round-trips past the deadline (shrinking the last step's budget to what remains) and answers with the same capped reply, never a 5xx.
Most of the system prompt is the same on every request: the base prompt
plus the HAE, combination-card, docs and category catalogues. The agent
loop re-sends all of it on every round-trip, and with CHAT_MAX_TURNS
at 12 a single question can transmit it a dozen times.
Gateways that support prompt caching serve a marked prefix from cache at roughly a tenth of the normal input price. So the system message is sent as ordered content blocks with cache breakpoints rather than one flat string:
| Segment | Contents | Breakpoint | Invalidated by |
|---|---|---|---|
| static | base prompt + HAE / combination-card / docs / category catalogues | yes | a deploy |
| instance | card list + reports catalogue | yes | adding, renaming or deleting a card or report |
| volatile | today's date + the card in focus | no | every request |
Caching is a prefix match: the gateway hashes everything up to a breakpoint and only serves a hit if every byte before it is identical. That is why the order matters and why the volatile blocks come last and carry no breakpoint of their own. Marking them would write a fresh cache entry on every request, which on a gateway that bills cache writes above uncached input is worse than not caching at all. Tool schemas are hashed ahead of the system message, which is fine: they are static.
A breakpoint caches everything up to itself, not just its own block, so the second one covers static+instance together. Creating a card misses on the second breakpoint and still hits on the first.
Voice mode keeps its output-format envelope in front of the static text, because that text says "Original system prompt follows". A voice turn therefore gets its own cache entry rather than sharing the text-mode one.
Two behaviours were measured against real gateways rather than assumed, and both are worth re-checking whenever the model behind the gateway changes:
- Nothing caches without the breakpoint. A control request carrying the same prefix with no breakpoint returned a zero hit rate on every attempt against one gateway, and an unreliable partial hit against another. Implicit caching is not something to count on.
- The breakpoint does not always bound the cache unit. One gateway
serves the cached prefix even when the volatile tail changes, which is
the intent. Another treats the whole system instruction as one cache
unit, so any tail change re-writes the prefix rather than reading it.
That costs the gap between a write and a read rather than breaking
anything. Growing
messagesbreaks neither.
Set CHAT_PROMPT_CACHE=0 to send the old flat string instead. The
blocks are standard OpenAI-compatible content parts, but a gateway that
rejects an array-valued content on the system role would fail every
chat request, and the flat form is byte-identical to what earlier
versions sent. Ordering is unaffected by the switch: it changes the
payload shape, not which blocks come first.
Cache effectiveness is visible in debug logs
(HEALTH_DEBUG=1) on each step, as cached= and cwrite= token
counts. usage=none means the gateway reported no usage at all, which
is different from reporting a zero hit rate.
POST /api/chat with stream: true in the body switches the response
to server-sent events, so the client can show live progress instead of
a spinner for the whole agent loop. The request is otherwise identical;
error statuses that fire before the stream opens (400, 503) stay plain
JSON, so clients should check the response Content-Type.
Events, in order of appearance:
| Event | Data | Meaning |
|---|---|---|
status |
{phase:'thinking'} |
a gateway round-trip started |
status |
{phase:'tool', tool, id?} |
a tool call is executing (id = manifest id when known) |
token |
{text} |
a fragment of the assistant's text, in order |
reset |
{} |
drop text streamed so far: it preceded tool calls and was not the answer |
reply |
same object the buffered mode returns (reply, speak?, followup?, capped?) |
the final payload |
error |
{error, status} |
the classified failure copy plus the status the buffered mode would have sent |
done |
{} |
terminator; always the last event |
Voice-mode turns (voiceMode: true) emit no token events: the model's
raw output is a JSON speak/display envelope nobody should watch being
typed. Status events still flow, and the reply event carries the same
{reply, speak} shape as the buffered path.
The gateway leg streams too (stream: true on the upstream call), and
the per-step timeout becomes an idle timeout there: a healthy long
generation keeps resetting it, a stalled one still trips it. A gateway
that ignores stream: true and answers buffered JSON is tolerated.
Comment heartbeats (: ping) are sent every 15s, and the response sets
X-Accel-Buffering: no so an nginx in front does not buffer the
stream. A client that disconnects mid-turn only mutes the events: the
loop runs to completion, matching the buffered path's semantics.
/api/conversations stores named transcripts in the per-instance
database (list by recency / create / fetch / rename / replace messages /
delete; hard caps of 100 conversations and 200 messages each).
POST /api/conversations/search takes { q } and answers the same
summaries the list does, filtered to conversations whose title or
message text contains q, plus a snippet of the matching line when
the hit was in the transcript. The needle rides in the body, not a query
string, because it is chat text and access logs record URLs. Matching is
a case-insensitive substring (an escaped literal, so . and * are
themselves) over a scan of every row, which the 100-conversation cap
keeps cheap enough to need no index.
When
POST /api/chat carries a conversationId, the server owns the
transcript:
- The request's
messagesare just the new turn. They are persisted before the loop runs, so a failed turn still shows the user's message when the conversation is reopened. - The loop is fed a window over the stored transcript: the newest messages that fit a ~24k-character budget (the newest always goes through). Per-turn gateway cost stops growing with conversation length; the model re-reads older state through its tools when it needs it.
- The shaped reply is appended after the turn, whether or not the
client is still connected. Voice replies persist
hasVoice; capped replies persistcapped, so a reloaded client can re-offer play and keep-going. - An untitled conversation gets a short model-generated title from an async side-call after the exchange completes (2-6 words, quotes stripped, 60-char cap). It never blocks or fails the turn; an unnamed conversation just stays unnamed until a later turn retries. The call rides the same gateway and counts against the normal chat allowance.
Without a conversationId the endpoint behaves exactly as before
(client-supplied transcript, nothing persisted server-side). The legacy
/api/chat/history endpoints remain until the client cutover.
A conversation turn is a server-side job that survives its client: iOS suspends a backgrounded tab and aborts its fetches, and without this an in-flight turn's reply had nowhere to go. Mechanics:
- One turn at a time per conversation. A concurrent
POST /api/chatfor the same conversation answers 409 before its message is persisted, so a retry after the running turn cannot double up the transcript. - Every event of a conversation turn is buffered with an id (streamed and buffered requests alike) and fanned out to any number of attached event streams.
GET /api/chat/turn/:conversationIdreattaches: buffered events replay fromLast-Event-ID(or?after=N), then the stream stays live untildone. 204 means there is nothing to attach to: read the conversation, where any completed reply is already persisted.- Completed turns linger for 30s so a client that missed
donecan still replay; after that the conversation is the durable copy. A replay that lost its head to the per-turn event cap (5000) starts with areset.
The intended client loop: on visibilitychange back to visible, hit
the reattach endpoint; on 204, refresh the conversation.
DELETE /api/chat/turn/:conversationId stops the running turn: the
loop halts at its next checkpoint (between round-trips or tool calls),
the user's message stays, no reply is persisted, the streamed response
ends with a stopped event (buffered callers get {stopped: true}),
and the one-turn lock releases. 404 when nothing is running.
The shipped widget uses all of this: turns are streamed conversation
requests (the send button becomes a stop button mid-turn), a legacy
chat/history.json transcript is folded into a conversation the first
time the new client loads, and the legacy endpoints remain only for
that import path.
Older deploys used CHAT_GATEWAY_HOST + CHAT_GATEWAY_PORT +
CHAT_GATEWAY_TLS + CHAT_GATEWAY_TOKEN + CHAT_GATEWAY_MODEL. These
still work; they're composed into the canonical CHAT_ENDPOINT_URL
internally. New installs should use the canonical names directly.
- It does not run tools directly. If your endpoint supports tool-use and you want the chat to use tools, that's the endpoint's responsibility.
- It does not embed model credentials beyond
CHAT_API_KEY. One bearer token per Klebb instance.
When an external agent (chat agent integration, cron job, mobile shortcut, etc.) needs to write to a card without going through the chat widget, it authenticates with a bearer token instead of a WebAuthn session cookie.
Enable it: set AGENT_API_TOKEN in the environment. Any request with
Authorization: Bearer <token> is treated as an authenticated agent.
Endpoints available to agents:
| Method | Path | Purpose |
|---|---|---|
GET |
/api/manifests |
List all cards |
POST |
/api/manifests |
Create a new card from a full manifest body |
GET |
/api/manifests/:id |
Fetch full manifest |
GET |
/api/manifests/:id/data |
Fetch just the data block |
POST |
/api/manifests/:id/data |
Replace the data block |
DELETE |
/api/manifests/:id |
Remove the card and its file |
POST |
/api/manifests/reorder |
Reassign meta.order across cards |
GET |
/api/views/:view |
List cards for a view |
GET |
/api/settings/cards |
List all cards with enable state |
POST |
/api/settings/cards/:id/enable |
Master enable |
POST |
/api/settings/cards/:id/disable |
Master disable |
POST /api/manifests/reorder
Authorization: Bearer $AGENT_API_TOKEN
Content-Type: application/json
{ "order": ["mood", "weight", "bp", "peptides"] }Writes sparse-numbered meta.order (100, 200, 300, …) to each listed
card. Unlisted cards keep their existing order. Any unknown id causes
a 404 with no writes performed. Returns { ok: true, updated: [...ids] }.
Idempotent — posting the same order twice is a no-op (file mtimes unchanged). Use this when the user says "move the mood card to the top" or similar.
To add a weight entry:
POST /api/manifests/weight/data
Authorization: Bearer <AGENT_API_TOKEN>
Content-Type: application/json
{
"data": [
{ "date": "2026-04-20", "kg": 85.5 },
{ "date": "2026-04-21", "kg": 86.0 }
]
}The POST replaces the entire data array. Agents are responsible for the upsert semantics (fetch, merge, write back).
The registry validates $schema and core meta.* fields, but does not
impose a schema on the data block. Whatever you write is what gets read
back. This keeps card authorship flexible but means agents must honour the
per-card convention (typically { date: "YYYY-MM-DD", ...fields } rows).
Agents can author brand new cards without filesystem access. POST /api/manifests takes a full manifest body and writes it to
$HEALTH_HOME/data/<meta.id>.json:
POST /api/manifests
Authorization: Bearer $AGENT_API_TOKEN
Content-Type: application/json
{
"$schema": "klebb.datafile.v1",
"meta": {
"id": "blood-pressure",
"label": "Blood Pressure",
"emoji": "🩺",
"view": { "enabled": true, "component": "list-card" },
"writeable": {
"fromWebapp": true, "todayAllowed": true, "pastAllowed": true,
"inputs": [
{"key":"systolic","label":"Systolic","type":"number","required":true},
{"key":"diastolic","label":"Diastolic","type":"number","required":true}
]
}
},
"description": "Home BP readings.",
"data": []
}| Status | Meaning |
|---|---|
| 201 | Created ({ok, id, source}) |
| 400 | Malformed: bad JSON, wrong $schema, missing meta.id/meta.label |
| 401 | No auth |
| 409 | meta.id already in use |
| 422 | meta.id fails the sanitiser (format / reserved / path escape) |
| 500 | Filesystem write failed |
meta.id must match /^[a-z0-9][a-z0-9._-]*$/, max 64 chars, and not
be one of the reserved names (_archive, _virtual, _meta,
auto-export, reports, index). Everything else is pass-through —
unknown renderer names are accepted on purpose; they render as a
placeholder card until a matching renderer ships. See
MANIFEST-SCHEMA.md for the full field reference.
Deletion mirrors the create path:
DELETE /api/manifests/blood-pressure
Authorization: Bearer $AGENT_API_TOKENReturns {ok, id} on success, 404 if the id is unknown. The file is
unlinked; any data it contained is gone. Prefer meta.enabled:false
if you only want to hide the card.
If a card has meta.enabled: false, writes to it still succeed (the data
is saved to disk). The user just won't SEE the card until they re-enable it.
This is intentional — you might want a background logging card that shows
up only when the user flips it on.
Here's the minimum viable agent:
import requests
TOKEN = "your-agent-token"
BASE = "https://health.example.com"
def log_weight(kg, date=None):
# Fetch current data
r = requests.get(f"{BASE}/api/manifests/weight/data",
headers={"Authorization": f"Bearer {TOKEN}"})
r.raise_for_status()
data = r.json()["data"]
# Upsert
date = date or datetime.date.today().isoformat()
data = [d for d in data if d.get("date") != date]
data.append({"date": date, "kg": kg})
data.sort(key=lambda d: d["date"])
# Write back
r = requests.post(f"{BASE}/api/manifests/weight/data",
headers={"Authorization": f"Bearer {TOKEN}"},
json={"data": data})
r.raise_for_status()Scale it up with card discovery (GET /api/manifests), schema inspection
(GET /api/manifests/:id), and whatever UX your agent provides.
A minimal chat agent integration wraps this API with a bearer token. The pattern is the same as any HTTP client integration:
- Skill reads
EDDZHEALTH_URL+EDDZHEALTH_TOKENfrom env - Intent dispatch: "log weight 85kg" →
POST /api/manifests/weight/data - Query dispatch: "what was yesterday's mood?" →
GET /api/manifests/mood/data→ filter → reply
You don't need any particular agent framework. Any HTTP client + your model of choice can drive this API.
AGENT_API_TOKENbypasses WebAuthn — treat it like a production secret.- HTTPS is strongly recommended; the session-cookie path explicitly sets
Secureon HTTPS origins. - The agent endpoint doesn't have a rate limit by default. Put one in front via nginx/Caddy if you expose it to the internet.
- Cards with
meta.writeable.pastAllowed: falseorfutureAllowed: falsewill reject out-of-window writes from the webapp UI, but the bearer-token agent API does NOT currently enforce these policies. Agents are trusted.