Skip to content

About

Async multi-agent LLM orchestrator with Elo routing, self-healing, and an OpenAI-compatible Proxy API. Built for Home Assistant, usable everywhere.

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

44 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

⚑ vromvrom-engine

Async multi-agent LLM orchestrator β€” hybrid routing, Elo scoring, self-healing

Vroom Vroom β€” running at full throttle 🏁

Python FastAPI asyncio License Tests CI Python CI & Security ToolMaker Validate

Built for Home Assistant Β· Works with any project

⚠️ Project Status: Active Development Personal tool shared with the community. Backend (DAG, async, Elo routing) is used daily. The React dashboard (ihm-v2/) is included β€” build it once, then gui_server.py serves it at /. Expect rough edges.

πŸ‡«πŸ‡· Version franΓ§aise disponible β†’ README.fr.md Β Β·Β  πŸ—ΊοΈ Which LLMs to use? β†’ STRATEGIES.md


Stop configuring. Start learning. vromvrom-engine routes your tasks to the right LLM automatically β€” using Elo scoring that adapts from your own usage history. Async, cost-aware, self-healing. Built on a Raspberry Pi budget.

🎯 What is it?

vromvrom-engine is a fully asynchronous (asyncio) multi-agent orchestration engine that coordinates multiple LLMs to solve complex tasks. It was born from the need to drive a home automation display (M5Stack Tab5) through Home Assistant β€” but its architecture is generic and reusable for any project.

πŸ’‘ Experiment with GPT-4o & Llama 70B for free The engine comes pre-configured to connect to the GitHub Models API, offering free access (subject to rate limits) to models like GPT-4o and Llama 70B for your tests using your standard GitHub account.

What makes it different

Feature Description
🧠 4-level hybrid routing Regex β†’ ML (sklearn) β†’ LLM β†’ Elo scoring β€” 0ms to 200ms depending on complexity
πŸ† Dynamic Elo scoring Each model gains/loses Elo points per task domain. Best model selected automatically
πŸ”„ Async parallel DAG Independent tasks run concurrently via asyncio.gather()
πŸ›‘οΈ Self-Healing Circuit breaker + exponential backoff retry + automatic provider fallback
πŸ’° Cost-aware routing Availability/budget cascade: falls back to the next configured provider when one is unavailable, rate-limited, or over budget β€” not a quality-based escalation
πŸ” Local hybrid RAG TF-IDF + BM25 + ChromaDB Embeddings fused via RRF (k=60) β€” zero cloud cost
πŸ‘οΈ HITL Human-In-The-Loop: pause/resume orchestration for human validation
πŸ“Š FastAPI dashboard Glassmorphism HTML/JS UI with real-time SSE, workflow editor, Elo charts
πŸ”Œ Drop-in Proxy API 100% OpenAI-compatible API (/v1/chat/completions). Plug Cursor, Cline, or Continue.dev directly into the engine
🏠 Native Home Assistant Built-in /api/ha routes and entity ingestion tools to let LLMs see and control your smart home
🌐 Distributed Swarm Dispatch tasks to remote workers (Raspberry Pi, VMs, etc.)
πŸ”Œ Plugin system Add custom agents via plugins/<name>/agent.py + plugin.json

πŸ—οΈ Architecture

vromvrom-engine/
β”œβ”€β”€ gui_server.py          # FastAPI entry point (lifespan + 15 routers)
β”œβ”€β”€ main.py                # CLI β€” direct launch without HTTP
β”‚
β”œβ”€β”€ core/                  # Engine core
β”‚   β”œβ”€β”€ engine.py          # Main orchestrator (DAG β†’ Agents)
β”‚   β”œβ”€β”€ llm_gateway.py     # Multi-provider gateway (18+ providers)
β”‚   β”œβ”€β”€ router.py          # Hybrid routing (fast-path + ML + LLM + Elo)
β”‚   β”œβ”€β”€ dag_runner.py      # Async parallel execution by stages
β”‚   β”œβ”€β”€ factory.py         # Agent instantiation (Planner/Executor/Reviewer)
β”‚   β”œβ”€β”€ state.py           # Thread-safe Pydantic GlobalState
β”‚   β”œβ”€β”€ checkpoint.py      # ACID snapshots (SQLite WAL)
β”‚   β”œβ”€β”€ healing.py         # Self-healing + retry
β”‚   β”œβ”€β”€ review_loop.py     # Reviewer β†’ Correction loop
β”‚   β”œβ”€β”€ elo_scorer.py      # Per-domain model Elo scoring
β”‚   β”œβ”€β”€ elo_router.py      # Routing type Elo scoring
β”‚   β”œβ”€β”€ circuit_breaker.py # Async circuit breaker (CLOSED/OPEN/HALF_OPEN)
β”‚   β”œβ”€β”€ hitl.py            # Human-In-The-Loop (asyncio.Event)
β”‚   β”œβ”€β”€ models_db.py       # SQLite SSOT β€” model catalog, pricing, quotas
β”‚   └── ...
β”‚
β”œβ”€β”€ agents/                # Specialized agents
β”‚   β”œβ”€β”€ planner.py         # Breaks task into a JSON DAG
β”‚   β”œβ”€β”€ executor.py        # Executes tasks (ReAct loop + tools)
β”‚   β”œβ”€β”€ reviewer.py        # Validates result quality
β”‚   └── tool_maker_agent.py  # Generates new Python tools automatically
β”‚
β”œβ”€β”€ memory/                # Semantic memory + RAG
β”‚   β”œβ”€β”€ rag.py             # Hybrid RAG (TF-IDF + BM25 + Embeddings + RRF)
β”‚   β”œβ”€β”€ facts.py           # Fact store (SQLite FTS5 BM25)
β”‚   β”œβ”€β”€ episodes.py        # Episodic memory (Jaccard similarity)
β”‚   β”œβ”€β”€ embeddings.py      # ChromaDB vector store
β”‚   └── skills.py          # Procedural memory (successful tool sequences)
β”‚
β”œβ”€β”€ tools/                 # Agent-usable tools
β”‚   β”œβ”€β”€ tool_registry.py   # Registry with per-tool timeouts
β”‚   └── sanitizer.py       # Secret masking (6 patterns)
β”‚
β”œβ”€β”€ api/routes/            # 15 FastAPI route modules
β”œβ”€β”€ services/              # Business logic decoupled from HTTP
β”œβ”€β”€ plugins/               # Custom plugins (dynamic loading)
β”œβ”€β”€ workflows/             # JSON workflow definitions (graphs)
β”œβ”€β”€ static/                # HTML/JS/CSS UI (glassmorphism, fallback if ihm-v2 is not built)
β”œβ”€β”€ ihm-v2/                # React/Vite/TS dashboard (Chat, HA, LLM registry, Prompt Studio…)
β”œβ”€β”€ tests/                 # 60 pytest files
└── docs/                  # Architecture documentation

Hybrid routing pipeline

User request
      β”‚
      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 1. Regex Fast-Path  β”‚ ──→ 0ms    (simple commands detected by pattern)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚ ambiguous
      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 2. ML Router sklearnβ”‚ ──→ 0ms    (local classifier, 75% confidence threshold)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚ confidence < 75%
      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 3. LLM Slow-Path    β”‚ ──→ ~200ms (Gemini Flash to resolve ambiguity)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚
      β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 4. Elo Scoring      β”‚ ──→ selects the best model for the task domain
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
      β”‚
      β–Ό
  Planner β†’ DAG β†’ Parallel Executor(s) β†’ Reviewer β†’ Response

βš™οΈ Quick Start

Requirements

  • Python 3.11+
  • At least one LLM API key (free Gemini tier works, see .env.example)

Setup

# 1. Clone the repo
git clone https://github.com/Axellum/vromvrom-engine.git
cd vromvrom-engine

# 2. Virtual environment
python -m venv .venv
# Windows:
.venv\Scripts\activate
# Linux/Mac:
source .venv/bin/activate

# 3. Install dependencies
pip install -r requirements.txt

# 4. Configure
cp .env.example .env
# Edit .env β€” at minimum set GEMINI_API_KEY (free at aistudio.google.com)

# 5. Engine config
cp config.example.json config.json
# Optional: adjust default models

# 6. Launch
#    (optional) React UI: cd ihm-v2 && npm install && npm run build
python gui_server.py
# β†’ Dashboard at http://localhost:8000

Minimal configuration (.env)

The engine works with a single free key:

# Free Gemini key: https://aistudio.google.com/apikey
GEMINI_API_KEY=AIza...

# Local API authentication key (generate a random value)
MOTEUR_API_KEY=change-me-with-secrets-token-urlsafe-32

All other providers (DeepSeek, Anthropic, Mistral...) are optional β€” the engine degrades gracefully.


πŸš€ Usage

Web Dashboard

cd ihm-v2 && npm install && npm run build   # β†’ ihm-v2/dist/
python gui_server.py                        # http://localhost:8000

If ihm-v2/dist is missing, the engine falls back to the legacy UI in static/ (also under /v1).

Dev with HMR: cd ihm-v2 && npm run dev (Vite :5173, proxies /api to :8000).

The React dashboard includes Chat (SSE + tools), Home Assistant / voice, LLM registry, Prompt Studio, observability, and config.

Via the REST API

import httpx

response = httpx.post(
    "http://localhost:8000/api/execute",
    json={"message": "Analyze this Python code and suggest improvements"},
    headers={"Authorization": "Bearer <MOTEUR_API_KEY>"}
)
print(response.json()["result"])

As a Drop-in Proxy for IDEs (Cursor, Cline, Continue)

vromvrom-engine exposes a standard OpenAI-compatible endpoint. You can point your favorite AI coding assistant to the engine to benefit from free routing and the circuit breaker:

  • Base URL: http://localhost:8000/v1
  • API Key: <MOTEUR_API_KEY> (from your .env)
  • Model: gemini-2.0-flash (or any model available in your routing)

🏠 Home Assistant Integration (Optional but Native)

vromvrom-engine ships with built-in tools (ha_entity_ingest.py) and routes (/api/ha/*) to interact with your Home Assistant instance out of the box.

Just add these to your .env:

HASS_URL=http://192.168.1.X:8123
HASS_TOKEN=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...

Your agents can now instantly read your sensors, check states, and control devices without writing any extra YAML!

Via the CLI

python main.py "Explain hexagonal architecture in 3 key points"

πŸ”Œ Supported LLM Providers

Provider Free tier Notes
Gemini (Google AI Studio) βœ… Recommended to start β€” generous free tier
GitHub Models βœ… GPT-4o, Llama 3.3 70B via GitHub token
DeepSeek πŸ’° Excellent price/quality ratio (~$0.14/M tokens)
Anthropic Claude πŸ’° Via API or Claude Code CLI (Pro subscription)
Mistral βœ… European GDPR-friendly models β€” free tier
OpenRouter πŸ’° Aggregator (200+ models)
LM Studio βœ… Local inference (Qwen, Llama, Mistral...)
Ollama βœ… Alternative local inference
MiniMax πŸ’° Built-in <think> reasoning
Cohere βœ… Free trial models (Command R / R+)
Cerebras βœ… 100% free β€” ultra-fast inference (GPT-OSS 120B, GLM 4.7)
Zhipu AI (GLM) πŸ’° 8 GLM models, strong in code/agentic tasks
xAI (Grok) πŸ’° Paid only

πŸ§ͺ Tests

# Install dev dependencies
pip install -r requirements-dev.txt

# Run all tests
pytest

# With coverage report
pytest --cov=. --cov-report=html

πŸ—ΊοΈ Which LLMs should I use?

See STRATEGIES.md for 4 ready-to-use configurations:

  • πŸ†“ Free & Local only β€” $0/month
  • πŸ‰ Chinese models β€” ~$5–15/month, near-frontier quality
  • ⚑ Claude API/CLI β€” near-zero extra cost with a Pro subscription
  • πŸš€ Full Power + cost-optimized β€” state of the art at ~10Γ— lower cost

🀝 Contributing

Contributions are welcome! See CONTRIBUTING.md.

Priority areas:

  • 🐳 Dockerization (minimal image + docker-compose)
  • πŸ“– API documentation (see docs/API_REFERENCE.md)
  • πŸ§ͺ Increased test coverage
  • 🌍 Support for additional LLM providers
  • 🎨 UI improvements

πŸ“„ License

MIT β€” See LICENSE.


πŸ™ Inspirations


Built with ❀️ for the home automation and open-source AI community

About

Async multi-agent LLM orchestrator with Elo routing, self-healing, and an OpenAI-compatible Proxy API. Built for Home Assistant, usable everywhere.

Topics

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages