Drop-in FastAPI proxy for llama.cpp, Ollama, vLLM and similar backends. Automatically prunes, summarizes, and extracts the most relevant context from large inputs using advanced strategies (readagent, rlm, embeddings) so your local models answer accurately without cloud services.
privacy embeddings session-management summarization locomo fastapi vector-search edge-ai vllm chromadb local-llm llm-inference retrieval-augmented-generation llm-proxy gguf token-compression context-compaction dynamic-context-pruning extractive-retrieval bm25-bm
-
Updated
Sep 4, 2026 - Python