Async multi-agent LLM orchestrator β hybrid routing, Elo scoring, self-healing
Vroom Vroom β running at full throttle π
Built for Home Assistant Β· Works with any project
β οΈ Project Status: Active Development Personal tool shared with the community. Backend (DAG, async, Elo routing) is used daily. The React dashboard (ihm-v2/) is included β build it once, thengui_server.pyserves it at/. Expect rough edges.
π«π· Version franΓ§aise disponible β README.fr.md Β Β·Β πΊοΈ Which LLMs to use? β STRATEGIES.md
Stop configuring. Start learning. vromvrom-engine routes your tasks to the right LLM automatically β using Elo scoring that adapts from your own usage history. Async, cost-aware, self-healing. Built on a Raspberry Pi budget.
vromvrom-engine is a fully asynchronous (asyncio) multi-agent orchestration engine that coordinates multiple LLMs to solve complex tasks. It was born from the need to drive a home automation display (M5Stack Tab5) through Home Assistant β but its architecture is generic and reusable for any project.
π‘ Experiment with GPT-4o & Llama 70B for free The engine comes pre-configured to connect to the GitHub Models API, offering free access (subject to rate limits) to models like GPT-4o and Llama 70B for your tests using your standard GitHub account.
| Feature | Description |
|---|---|
| π§ 4-level hybrid routing | Regex β ML (sklearn) β LLM β Elo scoring β 0ms to 200ms depending on complexity |
| π Dynamic Elo scoring | Each model gains/loses Elo points per task domain. Best model selected automatically |
| π Async parallel DAG | Independent tasks run concurrently via asyncio.gather() |
| π‘οΈ Self-Healing | Circuit breaker + exponential backoff retry + automatic provider fallback |
| π° Cost-aware routing | Availability/budget cascade: falls back to the next configured provider when one is unavailable, rate-limited, or over budget β not a quality-based escalation |
| π Local hybrid RAG | TF-IDF + BM25 + ChromaDB Embeddings fused via RRF (k=60) β zero cloud cost |
| ποΈ HITL | Human-In-The-Loop: pause/resume orchestration for human validation |
| π FastAPI dashboard | Glassmorphism HTML/JS UI with real-time SSE, workflow editor, Elo charts |
| π Drop-in Proxy API | 100% OpenAI-compatible API (/v1/chat/completions). Plug Cursor, Cline, or Continue.dev directly into the engine |
| π Native Home Assistant | Built-in /api/ha routes and entity ingestion tools to let LLMs see and control your smart home |
| π Distributed Swarm | Dispatch tasks to remote workers (Raspberry Pi, VMs, etc.) |
| π Plugin system | Add custom agents via plugins/<name>/agent.py + plugin.json |
vromvrom-engine/
βββ gui_server.py # FastAPI entry point (lifespan + 15 routers)
βββ main.py # CLI β direct launch without HTTP
β
βββ core/ # Engine core
β βββ engine.py # Main orchestrator (DAG β Agents)
β βββ llm_gateway.py # Multi-provider gateway (18+ providers)
β βββ router.py # Hybrid routing (fast-path + ML + LLM + Elo)
β βββ dag_runner.py # Async parallel execution by stages
β βββ factory.py # Agent instantiation (Planner/Executor/Reviewer)
β βββ state.py # Thread-safe Pydantic GlobalState
β βββ checkpoint.py # ACID snapshots (SQLite WAL)
β βββ healing.py # Self-healing + retry
β βββ review_loop.py # Reviewer β Correction loop
β βββ elo_scorer.py # Per-domain model Elo scoring
β βββ elo_router.py # Routing type Elo scoring
β βββ circuit_breaker.py # Async circuit breaker (CLOSED/OPEN/HALF_OPEN)
β βββ hitl.py # Human-In-The-Loop (asyncio.Event)
β βββ models_db.py # SQLite SSOT β model catalog, pricing, quotas
β βββ ...
β
βββ agents/ # Specialized agents
β βββ planner.py # Breaks task into a JSON DAG
β βββ executor.py # Executes tasks (ReAct loop + tools)
β βββ reviewer.py # Validates result quality
β βββ tool_maker_agent.py # Generates new Python tools automatically
β
βββ memory/ # Semantic memory + RAG
β βββ rag.py # Hybrid RAG (TF-IDF + BM25 + Embeddings + RRF)
β βββ facts.py # Fact store (SQLite FTS5 BM25)
β βββ episodes.py # Episodic memory (Jaccard similarity)
β βββ embeddings.py # ChromaDB vector store
β βββ skills.py # Procedural memory (successful tool sequences)
β
βββ tools/ # Agent-usable tools
β βββ tool_registry.py # Registry with per-tool timeouts
β βββ sanitizer.py # Secret masking (6 patterns)
β
βββ api/routes/ # 15 FastAPI route modules
βββ services/ # Business logic decoupled from HTTP
βββ plugins/ # Custom plugins (dynamic loading)
βββ workflows/ # JSON workflow definitions (graphs)
βββ static/ # HTML/JS/CSS UI (glassmorphism, fallback if ihm-v2 is not built)
βββ ihm-v2/ # React/Vite/TS dashboard (Chat, HA, LLM registry, Prompt Studioβ¦)
βββ tests/ # 60 pytest files
βββ docs/ # Architecture documentation
User request
β
βΌ
βββββββββββββββββββββββ
β 1. Regex Fast-Path β βββ 0ms (simple commands detected by pattern)
βββββββββββββββββββββββ
β ambiguous
βΌ
βββββββββββββββββββββββ
β 2. ML Router sklearnβ βββ 0ms (local classifier, 75% confidence threshold)
βββββββββββββββββββββββ
β confidence < 75%
βΌ
βββββββββββββββββββββββ
β 3. LLM Slow-Path β βββ ~200ms (Gemini Flash to resolve ambiguity)
βββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββ
β 4. Elo Scoring β βββ selects the best model for the task domain
βββββββββββββββββββββββ
β
βΌ
Planner β DAG β Parallel Executor(s) β Reviewer β Response
- Python 3.11+
- At least one LLM API key (free Gemini tier works, see
.env.example)
# 1. Clone the repo
git clone https://github.com/Axellum/vromvrom-engine.git
cd vromvrom-engine
# 2. Virtual environment
python -m venv .venv
# Windows:
.venv\Scripts\activate
# Linux/Mac:
source .venv/bin/activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Configure
cp .env.example .env
# Edit .env β at minimum set GEMINI_API_KEY (free at aistudio.google.com)
# 5. Engine config
cp config.example.json config.json
# Optional: adjust default models
# 6. Launch
# (optional) React UI: cd ihm-v2 && npm install && npm run build
python gui_server.py
# β Dashboard at http://localhost:8000The engine works with a single free key:
# Free Gemini key: https://aistudio.google.com/apikey
GEMINI_API_KEY=AIza...
# Local API authentication key (generate a random value)
MOTEUR_API_KEY=change-me-with-secrets-token-urlsafe-32All other providers (DeepSeek, Anthropic, Mistral...) are optional β the engine degrades gracefully.
cd ihm-v2 && npm install && npm run build # β ihm-v2/dist/
python gui_server.py # http://localhost:8000If ihm-v2/dist is missing, the engine falls back to the legacy UI in static/ (also under /v1).
Dev with HMR: cd ihm-v2 && npm run dev (Vite :5173, proxies /api to :8000).
The React dashboard includes Chat (SSE + tools), Home Assistant / voice, LLM registry, Prompt Studio, observability, and config.
import httpx
response = httpx.post(
"http://localhost:8000/api/execute",
json={"message": "Analyze this Python code and suggest improvements"},
headers={"Authorization": "Bearer <MOTEUR_API_KEY>"}
)
print(response.json()["result"])vromvrom-engine exposes a standard OpenAI-compatible endpoint. You can point your favorite AI coding assistant to the engine to benefit from free routing and the circuit breaker:
- Base URL:
http://localhost:8000/v1 - API Key:
<MOTEUR_API_KEY>(from your.env) - Model:
gemini-2.0-flash(or any model available in your routing)
vromvrom-engine ships with built-in tools (ha_entity_ingest.py) and routes (/api/ha/*) to interact with your Home Assistant instance out of the box.
Just add these to your .env:
HASS_URL=http://192.168.1.X:8123
HASS_TOKEN=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...Your agents can now instantly read your sensors, check states, and control devices without writing any extra YAML!
python main.py "Explain hexagonal architecture in 3 key points"| Provider | Free tier | Notes |
|---|---|---|
| Gemini (Google AI Studio) | β | Recommended to start β generous free tier |
| GitHub Models | β | GPT-4o, Llama 3.3 70B via GitHub token |
| DeepSeek | π° | Excellent price/quality ratio (~$0.14/M tokens) |
| Anthropic Claude | π° | Via API or Claude Code CLI (Pro subscription) |
| Mistral | β | European GDPR-friendly models β free tier |
| OpenRouter | π° | Aggregator (200+ models) |
| LM Studio | β | Local inference (Qwen, Llama, Mistral...) |
| Ollama | β | Alternative local inference |
| MiniMax | π° | Built-in <think> reasoning |
| Cohere | β | Free trial models (Command R / R+) |
| Cerebras | β | 100% free β ultra-fast inference (GPT-OSS 120B, GLM 4.7) |
| Zhipu AI (GLM) | π° | 8 GLM models, strong in code/agentic tasks |
| xAI (Grok) | π° | Paid only |
# Install dev dependencies
pip install -r requirements-dev.txt
# Run all tests
pytest
# With coverage report
pytest --cov=. --cov-report=htmlSee STRATEGIES.md for 4 ready-to-use configurations:
- π Free & Local only β $0/month
- π Chinese models β ~$5β15/month, near-frontier quality
- β‘ Claude API/CLI β near-zero extra cost with a Pro subscription
- π Full Power + cost-optimized β state of the art at ~10Γ lower cost
Contributions are welcome! See CONTRIBUTING.md.
Priority areas:
- π³ Dockerization (minimal image + docker-compose)
- π API documentation (see docs/API_REFERENCE.md)
- π§ͺ Increased test coverage
- π Support for additional LLM providers
- π¨ UI improvements
MIT β See LICENSE.
- LangChain β orchestration patterns
- CrewAI β multi-agent coordination
- Home Assistant β home automation ecosystem
- FastAPI β API framework