Skip to content

Latest commit

 

History

History
314 lines (244 loc) · 14.7 KB

File metadata and controls

314 lines (244 loc) · 14.7 KB

English · 简体中文

API reference

← Back to README · Documentation index

Any OpenAI-compatible client works (Anthropic / Claude clients too — see Anthropic / Claude clients). Base URL http://localhost:3001/v1, unified key from the dashboard's Keys page. An interactive OpenAPI viewer covering every proxy endpoint is served at GET /v1/docs; the spec itself lives at GET /v1/openapi.json.

Chat completions

Python

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

resp = client.chat.completions.create(
    model="auto",  # let the router pick; or specify e.g. "gemini-2.5-flash"
    messages=[{"role": "user", "content": "Summarise the fall of Rome in one sentence."}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))

curl

curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto",
    "messages": [{"role": "user", "content": "hi"}]
  }'

Routing strategies (auto:*)

Plain auto follows your active fallback chain. Add a suffix to steer a single request instead — no dashboard changes needed:

  • auto:smart — favor the highest-intelligence models
  • auto:fast — favor measured speed (throughput and time-to-first-byte)
  • auto:cheap — budget-leaning; currently the same blend as balanced (everything in the pool is already free)
  • auto:reliable — favor recent success rate
  • auto:balanced — the default blend (reliability first, speed and intelligence split the rest)

These rank every enabled model, ignoring your chain order. Common synonyms resolve too (auto:fastest, auto:speed, auto:smartest, auto:cheapest, auto:budget, …), and the whole model string is case-insensitive.

curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto:fast",
    "messages": [{"role": "user", "content": "hi"}]
  }'

auto:<profile-name> routes through a named profile's chain instead of the active one, so different tools can use different chains through the same key:

curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "auto:coding",
    "messages": [{"role": "user", "content": "Write a binary search in Rust."}]
  }'

An unknown profile name returns a clear 400 rather than silently falling back. Profiles are named fallback chains (see Features) — create and switch them from the dashboard; whichever is active is what plain auto uses.

Streaming

stream = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "Stream me a haiku about SQLite."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Tool calling

Pass OpenAI-style tools and tool_choice; the assistant response round-trips back through the proxy exactly like the OpenAI API. Multi-step flows (assistant tool_callstool role follow-up → final answer) work across every provider the router can reach.

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city.",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

# 1. Model asks for a tool call
first = client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": "What's the weather in Karachi?"}],
    tools=tools,
    tool_choice="required",
)
call = first.choices[0].message.tool_calls[0]

# 2. You execute the tool, feed the result back
final = client.chat.completions.create(
    model="auto",
    messages=[
        {"role": "user", "content": "What's the weather in Karachi?"},
        first.choices[0].message,
        {"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 32, "cond": "sunny"}'},
    ],
    tools=tools,
)
print(final.choices[0].message.content)

Works with stream=True as well — you'll get delta.tool_calls chunks followed by a finish_reason: "tool_calls" close. Under the hood, OpenAI-compatible providers (Groq, Cerebras, Mistral, OpenRouter, GitHub Models, HuggingFace, Cloudflare, Cohere compat) get the request passed through; Gemini requests get translated into Google's functionDeclarations / functionResponse shape and the response is translated back.

Gemini Google Search grounding

Google's models can ground their answers in live Google Search results. Since the OpenAI wire format has no way to express that, request a tool named google_search and the Google provider translates it into Gemini's native grounding tool. It can be sent on its own or alongside your normal function tools.

resp = client.chat.completions.create(
    model="gemini-2.5-flash",  # pin a Google model so the request routes there
    messages=[{"role": "user", "content": "Who won the F1 race this weekend?"}],
    tools=[{"type": "function", "function": {"name": "google_search", "parameters": {}}}],
)
print(resp.choices[0].message.content)

Native Gemini API

Gemini SDKs and Gemini CLI can use the native /v1beta surface:

  • GET /v1beta/models
  • GET /v1beta/models/{model}
  • POST /v1beta/models/{model}:generateContent
  • POST /v1beta/models/{model}:streamGenerateContent (?alt=sse for Gemini CLI)
  • POST /v1beta/models/{model}:countTokens
curl "http://localhost:3001/v1beta/models/gemini-2.5-flash:generateContent" \
  -H "x-goog-api-key: freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{"contents":[{"role":"user","parts":[{"text":"hello"}]}]}'

Gemini model names like gemini-2.5-flash resolve through the Gemini family map (Keys → Gemini model mapping) rather than the catalog directly, so they route to Auto or whichever catalog model each family is pinned to. Catalog ids work verbatim too.

contents text, inline data, function calls/responses, system instructions, function declarations, structured JSON output, generation controls, and thinking budgets translate into the same internal chat/fallback pipeline. Bearer auth works too. Gemini's ?key= query parameter is accepted only below /v1beta; prefer a header because URL credentials leak into history and logs.

Ollama emulation

The opt-in Ollama surface implements tags, chat, generate, show, version, embed, and legacy embeddings under /api/*. Streaming is Ollama-compatible NDJSON. It defaults to off; choose open-loopback or key-required on Keys → Agents. Open-loopback checks the direct socket peer, so desktop LAN access cannot silently turn it into an unauthenticated LAN endpoint.

The dashboard also owns /api/embeddings. A request with a valid dashboard session continues to the dashboard handler; all other requests at that exact path are treated as Ollama legacy embeddings and follow the emulation policy.

Revocable URL tokens

Headerless clients can mirror models, chat completions, Responses, and Ollama-style chat/tags under /v1/t/{token}/…. These tokens are random, stored only as hashes, separately revocable, and are not the unified key. Create/revoke them on Keys → Agents.

Treat them as sensitive anyway: URLs are routinely retained by shell history, reverse proxies, browser history, and telemetry. Revocation is immediate.

Vision / image input

Send images with the standard OpenAI image_url content blocks (base64 data: URLs or http(s) URLs). When a request contains an image, the router restricts itself to vision-capable models and ignores text-only ones. Vision models are tagged with a Vision badge on the Fallback Chain page; the current set includes Gemini (2.5 / 3.x), Llama 4 Scout/Maverick (Groq, NVIDIA), GLM-4.6V Flash (Z.ai), Nemotron Nano 12B VL (OpenRouter), and GitHub's GPT-4o / GPT-4.1.

resp = client.chat.completions.create(
    model="auto",  # auto-routes to a vision model
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What's in this image?"},
            {"type": "image_url", "image_url": {"url": "data:image/png;base64,<...>"}},
        ],
    }],
)
print(resp.choices[0].message.content)

If no vision-capable model is enabled in your Fallback Chain, an image request returns a clear 422 (code: "no_vision_model") rather than silently dropping the image. (Image input on /v1/responses isn't supported yet — use /v1/chat/completions.)

Images & text-to-speech

POST /v1/images/generations and POST /v1/audio/speech route across the providers that serve media models, including custom OpenAI-compatible media endpoints. Browse and toggle them on the dashboard's Models → Image / Audio tabs.

Fusion (multi-model synthesis)

Request the virtual fusion model and the router fans your prompt out to a panel of diverse free models in parallel, then a judge model synthesizes one answer from the drafts. Panel, judge, and strategy are configurable on the dashboard's Fusion page or per request via the fusion field; each sub-call goes through normal routing, quotas, and analytics.

Response headers

Every response carries an X-Routed-Via: <platform>/<model> header so you can see which provider actually served each call. If a request fell over between providers, you'll also see X-Fallback-Attempts: N.

HTTP headers only carry printable ASCII, so a model id with characters outside that range (a Chinese name from a relay catalog, for example) is percent-encoded in the header — run the value through decodeURIComponent (or urllib.parse.unquote) to read it back.

The opt-in response cache can be toggled per request with X-FreeLLM-Cache: on|off — an exact-match in-memory LRU for identical non-streaming requests (canonical SHA-256 keys over the full request, TTL and temperature gates, saved-token stats on the dashboard). Off by default; cache hits consume zero provider quota.

When prompt compression is enabled, X-FreeLLM-Compress: off|on|lossless|standard|aggressive can disable or lower the configured mode for one request. It cannot raise the operator's configured mode. The response reports the effective mode and estimated savings, for example X-FreeLLM-Compress: standard; saved~=1840.

Embeddings

/v1/embeddings is OpenAI-compatible, with one deliberate difference from chat routing: failover never crosses models. Vectors from different models live in incompatible spaces — silently switching models would corrupt any vector store built on top of the proxy. So embeddings route by family (one model identity + dimension), and failover only walks the providers serving that same family.

resp = client.embeddings.create(
    model="auto",          # default family; or a family name like "bge-m3"
    input=["the quick brown fox", "pack my box with five dozen liquor jugs"],
)
print(len(resp.data), "vectors of", len(resp.data[0].embedding), "dims")
curl http://localhost:3001/v1/embeddings \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{"model": "auto", "input": "hello world"}'

model accepts auto (the configured default family), a family name, or a provider-specific model id (which resolves to its family). Available families:

Family (model) Dims Providers (failover order)
gemini-embedding-001 (default) 3072 Google
text-embedding-3-large 3072 GitHub Models
text-embedding-3-small 1536 GitHub Models
embed-v4.0 1536 Cohere
bge-m3 1024 Cloudflare → Hugging Face
qwen3-embedding-0.6b 1024 Cloudflare
nv-embedqa-e5-v5 1024 NVIDIA
llama-nemotron-embed-1b-v2 2048 NVIDIA
llama-nemotron-embed-vl-1b-v2 2048 NVIDIA → OpenRouter
embeddinggemma-300m 768 Cloudflare

The default family, per-provider toggles, and priorities live on the dashboard's Models → Embeddings page. Pick your family once and stick with it for a given vector store — that's the whole point of the family model.

Anthropic / Claude clients

FreeLLMAPI also speaks Anthropic's Messages API, so anything built for Claude — including Claude Code and the official Anthropic SDKs — can run against your free pool. Point the client at your server's origin (Anthropic clients append /v1/messages themselves) and authenticate with your unified key. Both x-api-key and Authorization: Bearer are accepted.

curl http://localhost:3001/v1/messages \
  -H "x-api-key: freellmapi-your-unified-key" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-5",
    "max_tokens": 256,
    "messages": [{"role": "user", "content": "hi"}]
  }'

Claude model names map to your free pool on the Keys → Anthropic tab: each family (default, opus, sonnet, haiku) routes to auto (the router picks a free model) or a model you pin. POST /v1/messages/count_tokens and a content-negotiated GET /v1/models (Anthropic shape when anthropic-version is sent) are implemented too. Streaming, system prompts, tool use, and image input all translate across the same router as the OpenAI endpoints.

Claude Code — point it at your server and start it:

On macOS / Linux (Bash):

export ANTHROPIC_BASE_URL=http://localhost:3001
export ANTHROPIC_AUTH_TOKEN=freellmapi-your-unified-key   # NOT ANTHROPIC_API_KEY
claude

On Windows (PowerShell):

$env:ANTHROPIC_BASE_URL="http://localhost:3001"
$env:ANTHROPIC_AUTH_TOKEN="freellmapi-your-unified-key"
claude

Use ANTHROPIC_AUTH_TOKEN (sent as a Bearer token), not ANTHROPIC_API_KEY — Claude Code treats a set ANTHROPIC_API_KEY as a conflicting first-party credential and refuses to start.