A Reachy Mini front end for Nous Research's hermes-agent, following ClawBody's architecture exactly: OpenAI Realtime is the embodied brain; hermes-agent is reached via a single ask_hermes tool. See reports/2026-05-03-hermes-body-reachy-mini-frontend.md for the architectural rationale.
| Decision | Choice | Why |
|---|---|---|
| Hermes host | This Mac, LAN-only for v1 | Fastest debug loop; defer remote access |
| Robot/sim split | Both — sim default, robot when motion matters | Iterate on prompts in sim; validate on hardware |
| Secrets | .env in project root, gitignored |
Standard, what ClawBody does |
| v1 scope | Full ClawBody parity | Voice + motion + camera + face tracking + head wobble + identity bootstrap + post-turn sync |
| Hermes backing model | OpenRouter | Try several models with one key |
| Identity | Dynamic from Hermes at startup | Robot speaks AS the Hermes agent; cross-channel memory |
| UI | CLI + Gradio (--gradio flag) |
Browser mic invaluable for sim debugging |
| Realtime model | gpt-realtime (or gpt-4o-realtime-preview fallback) |
Only realistic sub-second voice option |
Reachy Mini mic → OpenAI Realtime (brain + tool dispatch) → either local robot tools OR ask_hermes → response → Realtime TTS → Reachy Mini speakers. Continuous behaviors (face tracking, head wobble) run in local threads, never network-bound.
hermes-body/
├── .env # gitignored — OPENAI_API_KEY, HERMES_*, etc.
├── .env.example # committed template
├── .gitignore
├── pyproject.toml
├── README.md
├── PLAN.md # this file
├── reports/
│ └── 2026-05-03-hermes-body-reachy-mini-frontend.md
└── src/hermes_body/
├── __init__.py
├── config.py # env loading + validation
├── main.py # HermesBodyApp + HermesBodyCore + main()
├── hermes_bridge.py # ask(), get_agent_context(), sync_turn()
├── openai_realtime.py # OpenAIRealtimeHandler (port of clawbody)
├── moves.py # MovementManager + HeadLookMove (lift)
├── camera_worker.py # CameraWorker (lift)
├── gradio_app.py # launch_gradio()
├── prompts.py # ROBOT_BODY_INSTRUCTIONS + FALLBACK_IDENTITY
├── audio/
│ ├── __init__.py
│ └── head_wobbler.py # HeadWobbler (lift)
├── vision/
│ ├── __init__.py
│ ├── mediapipe_tracker.py # HeadTracker (MediaPipe variant, lift)
│ └── yolo_head_tracker.py # HeadTracker (YOLO variant, lift)
└── tools/
├── __init__.py
└── core_tools.py # TOOL_SPECS, dispatch_tool_call, _handle_*
Goal: all three pieces (Mac venv, Reachy sim, hermes gateway) running and reachable.
- Mac venv:
cd /Users/wschenk/The-Focus-AI/hermes-body python3.11 -m venv .venv && source .venv/bin/activate pip install --upgrade pip
- Install Reachy Mini SDK + simulator:
pip install "reachy-mini[mujoco]" - Smoke test the simulator:
mjpython -m reachy_mini.daemon.app.main --sim # 3D window should appear; dashboard at http://localhost:8000 - Install hermes-agent on Mac:
curl -fsSL https://raw.githubusercontent.com/NousResearch/hermes-agent/main/scripts/install.sh | bash source ~/.zshrc hermes setup # walk through; pick OpenRouter as provider, paste OPENROUTER_API_KEY
- Enable the API server:
cat >> ~/.hermes/.env <<'EOF' API_SERVER_ENABLED=true API_SERVER_KEY=hb-dev-$(openssl rand -hex 16) API_SERVER_HOST=0.0.0.0 # so the robot can reach it later API_SERVER_PORT=8642 EOF
- Start gateway and verify:
hermes gateway # In another terminal: curl http://localhost:8642/v1/chat/completions \ -H "Authorization: Bearer <key from .env>" \ -H "Content-Type: application/json" \ -d '{"model":"hermes-agent","messages":[{"role":"user","content":"hi"}]}'
Done when: sim window opens, hermes gateway prints [API Server] listening on http://0.0.0.0:8642, and the curl returns a 200 with a chat response.
Goal: importable Python package with the Reachy Mini App entry point registered.
.gitignore:.venv/ __pycache__/ *.egg-info/ .env ~/.hermes/images/.env.example:# OpenAI Realtime (the brain) OPENAI_API_KEY=sk-... OPENAI_REALTIME_MODEL=gpt-realtime OPENAI_VOICE=cedar # Hermes (the knowledge tool) HERMES_BASE_URL=http://localhost:8642/v1 HERMES_API_KEY=hb-dev-... HERMES_MODEL=hermes-agent # Robot/dev ROBOT_NAME= ENABLE_FACE_TRACKING=true HEAD_TRACKER_TYPE=mediapipe
pyproject.toml: the load-bearing config.[build-system] requires = ["setuptools>=61.0", "wheel"] build-backend = "setuptools.build_meta" [project] name = "hermes-body" version = "0.1.0" description = "Reachy Mini front end for Nous Research's hermes-agent" requires-python = ">=3.11" dependencies = [ "openai>=1.50.0", "fastrtc>=0.0.17", "numpy", "scipy", "python-dotenv", "httpx>=0.27", "websockets>=12.0", "opencv-python-headless", "gradio>=4.0", "mediapipe>=0.10.14", ] [project.scripts] hermes-body = "hermes_body.main:main" [project.entry-points."reachy_mini_apps"] hermes-body = "hermes_body.main:HermesBodyApp" [tool.setuptools.packages.find] where = ["src"]
- Stub modules: create empty
src/hermes_body/{__init__,config,main,hermes_bridge,openai_realtime,prompts,moves,camera_worker,gradio_app}.pyplusaudio/__init__.py,vision/__init__.py,tools/__init__.py, and a placeholderHermesBodyAppclass inmain.pyso install succeeds. config.py: load.env, expose typed config object,validate()returns a list of error strings.- Install editable + smoke test:
pip install -e . python -c "from hermes_body.main import HermesBodyApp; print(HermesBodyApp)"
Done when: pip install -e . succeeds, hermes-body --help works, and a fresh python -c "import hermes_body" doesn't error.
Goal: working ask() / get_agent_context() / sync_turn() against the running gateway.
hermes_bridge.py: implementHermesResponsedataclass andHermesBridgeclass (see Pattern 3 in the report). Methods:ask(query, *, image_b64=None, system_context=None) -> HermesResponseget_agent_context() -> str | None(with the verbatim ClawBody prompt for the identity blob)sync_turn(user_msg, assistant_msg) -> None(fire-and-forget OK)
- Image-attached
ask: whenimage_b64is provided, build the OpenAI vision content array ([{"type":"text",...},{"type":"image_url",...}]). - Error handling: distinguish
httpx.TimeoutException,HTTPStatusError, generic exceptions. Always return aHermesResponse— never raise. - Smoke script (
tests/smoke_hermes.py, throwaway):import asyncio from hermes_body.hermes_bridge import HermesBridge async def main(): h = HermesBridge() r = await h.ask("What's 2+2?") print("ask:", r) ctx = await h.get_agent_context() print("identity:", ctx[:200] if ctx else None) await h.sync_turn("hello", "hi there") print("sync ok") asyncio.run(main())
Done when: smoke script returns sensible answers from your running hermes gateway.
Goal: all the local non-network code copy-pasted in, imports renamed, runnable in isolation.
- Clone ClawBody (already at
/tmp/clawbody-srcfrom earlier research; re-clone if gone):git clone https://github.com/tomrikert/clawbody.git /tmp/clawbody-src. - Lift verbatim with rename:
src/reachy_mini_openclaw/moves.py→src/hermes_body/moves.pysrc/reachy_mini_openclaw/audio/head_wobbler.py→src/hermes_body/audio/head_wobbler.pysrc/reachy_mini_openclaw/camera_worker.py→src/hermes_body/camera_worker.pysrc/reachy_mini_openclaw/vision/mediapipe_tracker.py→src/hermes_body/vision/mediapipe_tracker.pysrc/reachy_mini_openclaw/vision/yolo_head_tracker.py→src/hermes_body/vision/yolo_head_tracker.py(optional — only if you want YOLO too)
- Sed the imports:
sed -i '' 's/reachy_mini_openclaw/hermes_body/g' src/hermes_body/**/*.py src/hermes_body/*.py. - Drop config dependencies: wherever lifted code does
from reachy_mini_openclaw.config import config, replace withfrom hermes_body.config import config. Add the relevant fields toconfig.py. - Sim test: with
reachy-mini-daemon --simrunning, write a 30-line script that connectsReachyMini(), instantiatesMovementManager, queues aHeadLookMove(direction="left"), and waits 2s. Watch the sim head turn.
Done when: the lifted modules import cleanly and a MovementManager test script makes the sim move.
Goal: TOOL_SPECS and dispatch_tool_call lifted, plus ask_hermes tool spec.
- Lift
tools/core_tools.pyfrom ClawBody verbatim (rename imports). Keep all handlers:_handle_look,_handle_camera,_handle_face_tracking,_handle_dance,_handle_emotion,_handle_stop_moves,_handle_idle. - Strip OpenClaw-specific stuff: the
_analyze_image_with_openaihelper stays (it uses OpenAI vision directly, no OpenClaw). The OpenClaw-fallback branch in_handle_cameragoes — replace with "no description available, just return the b64 image" so the model can still see the pixels via Realtime. ToolDependenciesdataclass: drop theopenclaw_bridgefield, addhermes_bridge: HermesBridge | None.- Add
ask_hermesspec (NOT incore_tools.py— keep it inopenai_realtime.pynext to other dynamic tool building, since whether it's offered depends on whetherhermes_bridgeis wired):{ "type": "function", "name": "ask_hermes", "description": ( "Query Hermes (your full agent brain) for things you don't know off the top " "of your head: current weather, calendar, news, web search, smart home " "control, persistent memory, custom skills. Do NOT use this for chitchat or " "things you already know." ), "parameters": { "type": "object", "properties": { "query": {"type": "string"}, "include_image": {"type": "boolean", "default": False}, }, "required": ["query"], }, } - Unit smoke: dispatch each tool with mock deps; verify
dispatch_tool_call("look", '{"direction":"left"}', deps)returns{"status":"success","direction":"left"}.
Done when: all robot tools dispatch correctly with a mocked ToolDependencies, and the ask_hermes spec is wired in _build_tools().
Goal: voice loop works end-to-end against the sim.
- Lift
openai_realtime.pyfrom ClawBody. Rename imports. Keep:OPENAI_SAMPLE_RATE = 24000ROBOT_BODY_INSTRUCTIONS(move toprompts.py, lightly editask_openclaw→ask_hermes)FALLBACK_IDENTITY(rewrite for Hermes-flavored generic identity)OpenAIRealtimeHandlerclass — full lift
- Substitutions:
openclaw_bridgefield →hermes_bridgefield (typeHermesBridge | None)_handle_openclaw_query→_handle_hermes_query(same shape — callself.hermes.askinstead ofself.openclaw_bridge.chat)_sync_to_openclaw→_sync_to_hermes(callself.hermes.sync_turn)_build_system_instructionscallsself.hermes.get_agent_context()instead ofself.openclaw_bridge.get_agent_context()- In
_build_tools(): appendask_hermesspec instead ofask_openclaw - In
_handle_tool_call(): dispatchask_hermesto_handle_hermes_query
- Reconnection loop: keep ClawBody's exponential-backoff reconnect in
start_up()verbatim — it handles WebSocket drops gracefully. - Logging: keep ClawBody's INFO-level
User: ...andAssistant: ...logs — they're how you'll debug everything.
Done when: the handler instantiates, reads OPENAI_API_KEY from config, and start_up() logs OpenAI Realtime session configured with N tools (where N = robot tools + 1 for ask_hermes).
Goal: hermes-body CLI works against the sim end-to-end.
HermesBodyCore: orchestrator; lift from ClawBody'sClawBodyCore:- Constructor takes
robot=None,external_stop_event=None - Wires up:
MovementManager,HeadWobbler,CameraWorker(with optionalHeadTracker),HermesBridge,OpenAIRealtimeHandler record_loop()/play_loop()— verbatim from ClawBodyrun()— enable motors, neutral pose, start movement/wobbler/camera/audio threads, gather tasksstop()— graceful shutdown of all subsystems
- Constructor takes
HermesBodyApp: thereachy_mini_appsentry-point class. Trivial wrapper that creates a new event loop and callsHermesBodyCore(robot=reachy_mini, external_stop_event=stop_event).run().main(): argparse with the same flags as ClawBody (--debug,--gradio,--robot-name,--no-camera,--no-hermes,--no-face-tracking,--head-tracker).- First end-to-end run:
Speak into your Mac mic. Say "hi". Hear a reply. Say "look left". Watch the sim turn.
# Terminal A: hermes gateway (already running from Phase 0) # Terminal B: simulator mjpython -m reachy_mini.daemon.app.main --sim # Terminal C: hermes-body hermes-body --debug
Done when: chitchat works, look fires the local tool, ask_hermes ("what's the weather?") fires the Hermes round trip, and response.done triggers sync_turn (visible in hermes gateway logs as a new message).
Goal: hermes-body --gradio launches a browser UI at localhost:7860 with mic + transcript.
- Lift
gradio_app.pyfrom ClawBody. It uses FastRTC'sStreamwith the sameOpenAIRealtimeHandler. Mostly a UI wrapper; rename imports. - Launch from
main.pywhen--gradiois set:from hermes_body.gradio_app import launch_gradio; launch_gradio(...). - Test: open
http://localhost:7860, click mic, speak, see transcript stream.
Done when: browser-based voice loop works in addition to the CLI mic loop.
Goal: confidence that everything works before touching the robot.
Run through this checklist with sim + hermes-body --debug + hermes gateway tail:
- Greeting: "Hi" → robot replies with personality from Hermes (identity blob is loaded). Verify the personality matches what
hermesCLI shows for the same agent. - Look tool: "Look to your left" →
_handle_lookfires, sim head turns left. - Dance tool: "Do a dance" →
_handle_dancefires, sim plays a move. - Emotion tool: "Show me you're curious" →
_handle_emotionfires, sim animates. - Camera tool: "What do you see?" →
_handle_camerafires; in sim, the camera returns the rendered scene; vision description comes back. - Face tracking toggle: "Stop tracking faces" →
_handle_face_tracking({"enabled":false})fires; verify the worker is paused. - ask_hermes (no image): "What's the date?" or "Search the web for X" →
_handle_hermes_queryfires; Hermes processes; reply spoken. - ask_hermes (with image): "Look at this and tell Hermes what you see" →
include_image=truepath; image arrives in Hermes logs. - Post-turn sync: check
hermes gatewaylogs after each turn —[ROBOT BODY SYNC]messages should appear. - Hermes down resilience: stop
hermes gateway, ask the robot something requiring Hermes; verify it speaks the tellable error instead of going silent. - Realtime reconnect: kill your network for 5s, restore; verify Realtime reconnects (check the exponential-backoff log).
Done when: every box above is checked.
Goal: hermes-body runs on the actual Reachy Mini.
- Get hermes-agent reachable from the robot. Find your Mac's LAN IP (
ipconfig getifaddr en0); update.envwithHERMES_BASE_URL=http://<mac-lan-ip>:8642/v1. Verify from the robot:If that fails: macOS firewall is probably blocking port 8642. System Settings → Network → Firewall → allow Python.ssh pollen@reachy-mini.local curl -s http://<mac-lan-ip>:8642/v1/models -H "Authorization: Bearer <key>" | head
- Push the code to the robot:
Or, if you've pushed to GitHub:
# Easiest: rsync from your dev machine rsync -av --exclude='.venv' --exclude='__pycache__' --exclude='.git' \ /Users/wschenk/The-Focus-AI/hermes-body/ \ pollen@reachy-mini.local:~/hermes-body/
ssh pollen@reachy-mini.local 'git clone <url> hermes-body'. - Install into apps_venv:
ssh pollen@reachy-mini.local cd ~/hermes-body /venvs/apps_venv/bin/pip install -e .
- Copy
.envto the robot:scp .env pollen@reachy-mini.local:~/hermes-body/.env. UpdateHERMES_BASE_URLto the Mac IP. - Verify dashboard discovery: open
http://reachy-mini.local:8000. The "hermes-body" app should appear in the list. Click Run. - Run from CLI for first test (easier to see logs):
Walk through the same Phase 8 checklist on real hardware.
ssh pollen@reachy-mini.local /venvs/apps_venv/bin/hermes-body --debug
Done when: the dashboard launches the app and the full Phase 8 checklist passes on the physical robot.
| Phase | Estimate |
|---|---|
| 0. Environment setup | 30 min |
| 1. Project skeleton | 30 min |
| 2. HermesBridge | 1 hour |
| 3. Lift movement/camera/vision | 2 hours |
| 4. Tools registry | 2 hours |
| 5. OpenAIRealtimeHandler | 3 hours |
| 6. HermesBodyCore + main | 2 hours |
| 7. Gradio UI | 1 hour |
| 8. Sim validation | 2 hours |
| 9. Robot deployment | 1 hour |
| Total | ~15 hours |
Realistically a long weekend.
- YOLO vs MediaPipe face tracker: v1 ships MediaPipe (lighter). Add YOLO option if accuracy matters.
- Local vision (SmolVLM2): ClawBody supports on-device vision. Skip for v1; revisit if Hermes-routed vision feels slow.
- Gradio scene picker: ClawBody has a
--scene minimalflag for the sim. Skip for v1. - Tailscale deployment: when you're ready to use the robot away from your Mac, do the tailnet setup.
- Profile per-user: Hermes supports
hermes -p <name>profiles for isolated identities. Useful if multiple people want their own robot brain.