Local-first policy and guidance layer for OpenClaw WhatsApp. It applies per-contact or per-group profiles and can generate drafts, controlled auto-replies, images, stickers, voice-note transcriptions, and audio replies.
The main flow is OpenClaw-only. Twilio support is included as an experimental webhook adapter for sandbox testing, not as the primary runtime.
- Per-contact and per-group profiles with
observe,draft, or gatedautoreplies. - WhatsApp voice-note transcription with direct API transcription or local
whisper.cpp. - Inbound image understanding/OCR, profile-gated with
tools.imageUnderstanding=true. - Structured weather lookup through Open-Meteo using WhatsApp shared locations, coordinates, or city/bairro text.
- Image generation, profile-gated with
tools.imageGeneration=true, delivered as WhatsApp media. - Native WhatsApp sticker generation, profile-gated with
tools.stickerGeneration=true, delivered through OpenClaw asasSticker=true. - Audio replies, profile-gated with
voice.reply.enabled=true, either on request or for every reply. - Optional
claude-proxyto back replies and inbound image understanding with theclaudeCLI (Claude Code subscription). - Optional
codex-proxyfor local Codex CLI-backed responses, local TTS, and local Whisper forwarding. - Cloudflare Workers AI image generation (
flux-1-schnell) for the claude backend, since theclaudeCLI cannot generate images.
- OpenClaw gateway: WhatsApp connection and delivery.
openclaw-control: local send endpoint for manual/test sends.openclaw-worker: inbound policy/profile worker.- Optional
claude-proxy: local OpenAI-compatible wrapper for theclaudeCLI (chat + vision). - Optional
codex-proxy: local OpenAI-compatible wrapper forcodex execor local transcription forwarding. - Optional
whisper-local: localwhisper.cpptranscription server for WhatsApp voice notes. - Optional Twilio webhook worker for sandbox testing.
All managed process logs and pid files are written under data/runtime/.
The public default is a direct OpenAI-compatible API key through RESPONDER_*. codex-proxy is optional for users who want to back replies with a local Codex CLI session.
Prerequisites:
- Node.js 20+ and npm.
- An OpenAI-compatible API key for the responder.
- OpenClaw CLI available on PATH, or
OPENCLAW_COMMANDset in.env. - Optional: Codex CLI installed and authenticated locally when using
CODEX_PROXY_ENABLED=true.
Windows:
npm install
copy .env.example .env
copy config\bot-policy.example.json config\bot-policy.local.jsonEdit .env and set:
RESPONDER_API_KEY=your-api-key
RESPONDER_MODEL=gpt-4o-mini
Optional media setup:
npm run media:install
npm run tts:install
npm run warmup:whisperThen opt profiles into the capabilities you want in config/bot-policy.local.json, for example tools.imageUnderstanding=true, tools.imageGeneration=true, tools.stickerGeneration=true, voice.enabled=true, or voice.reply.enabled=true. See Guidance profiles and Codex proxy for the provider-specific environment variables.
Then start:
npm run warmup
npm run warmup:statusLinux, Orange Pi, or a small VPS:
npm install
cp .env.example .env
cp config/bot-policy.example.json config/bot-policy.local.jsonEdit .env as above, then start:
npm run warmup:linux
npm run warmup:statusPair WhatsApp in OpenClaw when prompted by the gateway. Keep .env, data/, and config/bot-policy.local.json private.
Use npm install, not npm ci. package-lock.json is intentionally not committed while this project is still changing quickly. Forks that run this long-term should commit their generated lockfile for reproducible installs.
npm run warmup
npm run warmup:statusOn Linux, use npm run warmup:linux instead of npm run warmup.
warmup and warmup:linux install/refresh the local OpenClaw dispatch plugin, repair the OpenClaw config for this project, and start:
- OpenClaw gateway on
127.0.0.1:18789 openclaw-controlon127.0.0.1:8788openclaw-workeron127.0.0.1:8790- optional
codex-proxyon127.0.0.1:8787whenCODEX_PROXY_ENABLED=true - optional
whisper-localon127.0.0.1:2022whenWHISPER_LOCAL_ENABLED=true
To stop only processes started by the warmup manager:
npm run warmup:stopOn Windows, after setup, you can also double-click start-chatbot.bat to run warmup and status.
To start the same stack automatically when your Windows user logs in:
npm run service:install
npm run service:statusYou can also double-click install-chatbot-service.bat. The scheduled task runs start-chatbot.ps1 -NoPause -LogToFile, so startup output is written to data/runtime/start-chatbot.log. To remove only the startup task:
npm run service:uninstallFor 24/7 Linux hosting, first validate warmup:linux, then move the same services to systemd or another restart manager. See Hosting.
observe: log policy decisions only.draft: generate replies but do not auto-send.auto: auto-reply only when both global mode and target policy allow it.
Configure profiles and targets in config/bot-policy.local.json. Start from config/bot-policy.example.json.
The example policy opens inbound visibility with allowContacts=["*"] and allowGroups=true, but keeps defaults.mode="observe". That means unknown chats are visible to the worker but do not generate replies unless you add a target or intentionally change the defaults.
Profiles default to showing WhatsApp's native typing indicator while an automatic reply is being generated. Disable it per profile with typing.enabled=false, or tune the refresh interval with typing.intervalMs.
Profiles can opt into structured weather lookup with tools.weather=true. The agent can plan a get_weather action, then the worker resolves it with Open-Meteo using, in order, WhatsApp shared-location coordinates from OpenClaw metadata, decimal coordinates in the message, or a city/bairro from the planned query. If no location is available, the responder asks for one instead of using web search or guessing.
Profiles can opt into inbound image understanding with tools.imageUnderstanding=true. When WhatsApp sends an image with a local mediaPath, the worker extracts OCR/visual context before the responder runs and stores the image briefly as a visual reference for that chat. Direct API mode uses IMAGE_UNDERSTANDING_*; in local testing, IMAGE_UNDERSTANDING_PROVIDER=codex-cli lets Codex read the local image path. If image understanding fails, the worker falls back to a short text explanation instead of silently ignoring the image.
Inbound messages are processed sequentially per WhatsApp conversation. The worker keeps a short multimodal log with text messages, voice transcripts, and recent inbound image references; each response receives that recent context. With tools.localRead=true, Codex can inspect local image paths when a later prompt depends on images sent earlier in the chat.
Before tool side effects, the worker asks the agent for a structured action plan. Planned actions include get_weather, generate_image, generate_sticker, and reply_audio; the worker still enforces profile opt-in, delivery gates, rate limits, and provider configuration before executing anything.
Profiles can opt into image generation with tools.imageGeneration=true. When the agent plans generate_image, the worker calls the configured image provider, saves the generated file under MEDIA_OUTPUT_DIR, and sends it through OpenClaw with openclaw message send --media. If the planned action sets useRecentImages=true, the worker sends up to MEDIA_REFERENCE_MAX_IMAGES recent inbound images to the edit/reference endpoint instead of generating from text alone. With IMAGE_GENERATOR_PROVIDER=cloudflare (the default when CLOUDFLARE_ACCOUNT_ID + CLOUDFLARE_API_TOKEN are set) generation runs on Cloudflare Workers AI (flux-1-schnell); with IMAGE_GENERATOR_PROVIDER=openai it uses an OpenAI-compatible endpoint (IMAGE_GENERATOR_API_KEY/OPENAI_API_KEY, or CODEX_PROXY_MEDIA_PROVIDER=codex-cli for local Codex generation). See Image generation.
Profiles can opt into native WhatsApp sticker generation with tools.stickerGeneration=true. When the agent plans generate_sticker, the worker uses the same image provider, optionally includes recent inbound images as references, asks for a flat #00ff00 chroma-key source image, removes that key with FFmpeg, cleans transparent pixels with Pillow, writes a 512x512 lossless WebP with exact alpha, and sends it through OpenClaw's WhatsApp upload-file gateway action as asSticker=true. Run npm run media:install for the Pillow dependency. npm run warmup reapplies the local OpenClaw WhatsApp sticker patch after plugin install/refresh; configure MEDIA_FFMPEG_COMMAND or reuse CODEX_PROXY_FFMPEG_COMMAND.
Profiles can opt into audio replies with voice.reply.enabled=true. Use voice.reply.mode="on_request" to let the agent plan reply_audio when the conversation calls for it, or voice.reply.mode="always" to deliver every generated text reply as an audio file. In direct mode, configure SPEECH_API_KEY or OPENAI_API_KEY; in Codex proxy openai mode, speech uses the same upstream media provider. In codex-cli media mode, speech can use local System.Speech, Edge TTS, or Piper through the repo-local TTS adapter and emit Ogg/Opus through the portable FFmpeg installed by warmup:whisper. If speech generation fails, the worker falls back to text.
Profiles can opt into retroactive replies with retroactiveReply.enabled=true. The worker scans recent OpenClaw history for configured auto-reply targets and answers the latest inbound message that has no later own reply; retroactiveReply.maxAgeHours defaults to 12.
Profiles can also opt into WhatsApp voice-note transcription with voice.enabled=true. Defaults keep voice disabled. Direct API mode can transcribe with the same provider credentials; local transcription uses codex-proxy plus a whisper.cpp server:
CODEX_PROXY_ENABLED=true
CODEX_PROXY_TRANSCRIBER_PROVIDER=local-whisper
WHISPER_LOCAL_ENABLED=true
WHISPER_LOCAL_MODEL=base
Run npm run warmup:whisper once to download the local binaries/model, or let npm run warmup start it when WHISPER_LOCAL_ENABLED=true.
See Voice notes for setup, testing, and troubleshooting.
npm run openclaw:send -- --target +15551234567 --message "hello" --verboseOr through the local control daemon:
curl -X POST http://127.0.0.1:8788/send -H "Content-Type: application/json" -d "{\"target\":\"+15551234567\",\"message\":\"hello\",\"verbose\":true}"Default direct API mode:
CODEX_PROXY_ENABLED=false
RESPONDER_BASE_URL=https://api.openai.com/v1
RESPONDER_API_KEY=your-api-key
RESPONDER_MODEL=gpt-4o-mini
All-Cloudflare backend (recommended for self-hosting — chat/vision/image/transcription on Cloudflare Workers AI, web search via Tavily, TTS via local edge-tts; no local model/CLI/proxy):
CLOUDFLARE_ACCOUNT_ID=...
CLOUDFLARE_API_TOKEN=...
TAVILY_API_KEY=...
RESPONDER_PROVIDER=cloudflare
IMAGE_GENERATOR_PROVIDER=cloudflare
TRANSCRIBER_PROVIDER=cloudflare
SPEECH_PROVIDER=local
This is the setup the Docker / cloud deployment ships (e.g. Oracle Cloud Always Free). Chat models cannot browse, so the worker runs a Tavily web_search planner action and injects results.
Optional Claude Code proxy mode (chat + inbound image understanding via the claude CLI):
CLAUDE_PROXY_ENABLED=true
RESPONDER_BASE_URL=http://127.0.0.1:8789/v1
RESPONDER_API_KEY=dev-local-change-me
RESPONDER_MODEL=sonnet
CLAUDE_PROXY_ENABLED wins over CODEX_PROXY_ENABLED. The claude CLI cannot generate images, transcribe, or do TTS, so on the claude backend image generation uses Cloudflare Workers AI (IMAGE_GENERATOR_PROVIDER=cloudflare) while transcription/TTS stay on their own providers (e.g. local Whisper / edge via codex-proxy). See Claude proxy.
Optional Codex CLI proxy mode:
CODEX_PROXY_ENABLED=true
RESPONDER_BASE_URL=http://127.0.0.1:8787/v1
RESPONDER_API_KEY=dev-local-change-me
RESPONDER_MODEL=gpt-5.4
When CODEX_PROXY_ENABLED=true and CODEX_PROXY_MEDIA_PROVIDER is enabled, image generation and speech defaults also point at http://127.0.0.1:8787/v1 unless IMAGE_GENERATOR_BASE_URL or SPEECH_BASE_URL is set explicitly. Use CODEX_PROXY_MEDIA_PROVIDER=openai for upstream pass-through or CODEX_PROXY_MEDIA_PROVIDER=codex-cli for local Codex image generation plus local TTS.
See Claude proxy and Codex proxy and responder provider.
Twilio support is a helper path for sandbox testing. It reuses the same policy profiles and responder.
npm run twilio:workerTwilio is not started by warmup; run it separately. Expose only http://127.0.0.1:8791/twilio/whatsapp through a tunnel and configure that URL as the Twilio inbound webhook. The Twilio worker disables the OpenClaw /openclaw/message route on that port. For tunneled testing, keep TWILIO_VALIDATE_SIGNATURE=true and set TWILIO_AUTH_TOKEN plus TWILIO_WEBHOOK_URL. Do not commit real Twilio credentials.