English | 简体中文
A GOST Rewriter HTTP plugin that converts bidirectionally between OpenAI Chat Completions, Anthropic Messages, and Google Gemini generateContent API formats. Designed for use with tools like Claude Code, Codex CLI, OpenCode, and other LLM clients that speak any of these protocols.
- llm-api-converter
Deployed as a GOST rewriter plugin, it intercepts HTTP request/response bodies in a forward proxy and transparently converts between the two wire formats:
Anthropic SDK client → GOST → llm-api-converter → OpenAI-compatible API
↕ (protocol conversion) ↕
Anthropic format OpenAI format
The converter auto-detects the input format using positive structural markers only — no negative exclusions, no hardcoded model prefixes. A SessionStore tracks client protocol across request/response pairs for correct bidirectional routing. A URI-based fallback handles minimal requests that lack distinguishing features.
# Build
go build -o llm-api-converter .
# Run standalone
./llm-api-converter --addr :8000 --model-map "claude-opus=deepseek-v4-pro:openai,claude-sonnet=deepseek-v4-flash,*=deepseek-v4-flash:openai"
# Run with GOST
gost -C gost.yamlRun the converter as a container alongside GOST. The published image is ginuerzh/llm-api-converter (multi-arch: amd64/arm64/arm v6/v7). Since the image ENTRYPOINT is the binary, command: supplies the CLI flags.
# docker-compose.yml
services:
llm-converter:
image: ginuerzh/llm-api-converter:latest
command:
- --addr
- :8000
- --model
- deepseek-v4-flash
- --model-map
- claude-opus=deepseek-v4-pro:openai,claude-sonnet=deepseek-v4-flash,*=deepseek-v4-flash:openai
ports:
- "8000:8000"
restart: unless-stopped
gost:
image: gogost/gost:latest
command: -C /etc/gost/gost.yaml
volumes:
- ./gost.yaml:/etc/gost/gost.yaml:ro
ports:
- "8787:8787"
depends_on:
- llm-converter
restart: unless-stoppedPoint the GOST rewriter plugin at the converter's container address:
# in gost.yaml
rewriters:
- name: llm-converter
plugin:
type: http
addr: http://llm-converter:8000/rewritedocker compose up -d
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claudeBuild the image locally instead of pulling:
docker build -t ginuerzh/llm-api-converter .
# or, with the multi-arch buildx workflow from .github/workflows/buildx.ymlThis setup lets Claude Code (Anthropic protocol) call DeepSeek models (OpenAI protocol) through the converter:
Claude Code → GOST (proxy) → llm-api-converter → opencode-go API → DeepSeek
1. Start the converter:
./llm-api-converter \
--addr :8000 \
--model deepseek-v4-flash \
--model-map "claude-opus=deepseek-v4-pro:openai,*=deepseek-v4-flash:openai"2. Configure GOST to intercept Anthropic API calls and forward them through the converter:
# gost.yaml
services:
- name: claude-code-proxy
addr: :8787
handler:
type: tcp
metadata:
sniffing: true
listener:
type: tcp
forwarder:
nodes:
- name: opencode-go
addr: opencode.ai:443
tls:
secure: true
serverName: opencode.ai
http:
host: opencode.ai
rewriteURL:
# Anthropic /v1/messages → OpenAI /v1/chat/completions
- match: /v1/messages
replacement: /zen/go/v1/chat/completions
requestHeader:
Authorization: "Bearer your-oc-apikey"
x-api-key: "your-oc-apikey"
rewriteRequestBody:
- rewriter: llm-converter
type: application/json
rewriteResponseBody:
- rewriter: llm-converter
type: "*"
rewriters:
- name: llm-converter
plugin:
type: http
addr: http://127.0.0.1:8000/rewrite3. Point Claude Code at the proxy:
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claudeAll Anthropic traffic from Claude Code is intercepted by GOST, converted to OpenAI Chat Completions format by the plugin, and forwarded to the opencode-go API for DeepSeek inference. Responses and SSE streams are converted back to Anthropic format transparently.
Model mapping notes:
claude-opus=deepseek-v4-pro:openai: Routes requests with model name starting withclaude-opusto DeepSeek V4 Pro, converting Anthropic→OpenAI*=deepseek-v4-flash:openai: Catch-all fallback for any unmatched model prefix, also converting to OpenAI- Downstream protocol override: Append
:openai,:anthropic, or:geminiafter the target to declare what format the backend speaks (prefix=target:protocol). Without a protocol suffix, the default is passthrough — the body passes through with only the model name rewritten, no format conversion. With:openai/:anthropic/:gemini, conversion runs when the incoming protocol differs from the declared one; when they match, only the model is rewritten. The override applies on both request and response paths via per-session client protocol tracking. Example:claude-opus=deepseek-v4-pro:openai— incoming Anthropic differs from:openai, so Anthropic→OpenAI conversion runs;claude-opus=deepseek-v4-pro:anthropic— incoming Anthropic matches, so only the model is rewritten. - Note:
:responsesis not a valid override (onlyopenai/anthropic/gemini); Responses API traffic is detected and routed via body markers and the session store, not the model map. Empty targets (e.g.claude-opus=:openai) are rejected at parse time.
Update the --model-map to match your opencode-go deployment's available models.
A complete multi-vendor LLM router setup using GOST + llm-api-converter — routes requests to different providers based on prompt content, with automatic protocol conversion. See the LLM Router blog post for the full guide.
This setup lets Codex CLI (OpenAI Responses API protocol) call DeepSeek models (OpenAI Chat Completions protocol) through the converter:
Codex CLI → GOST (proxy) → llm-api-converter → opencode-go API → DeepSeek
Codex CLI sends Responses API format (POST /v1/responses with {model, input, ...}); the converter translates to Chat Completions format (POST /v1/chat/completions with {model, messages, ...}) for opencode-go, and reverses the response on the way back.
1. Start the converter:
./llm-api-converter \
--addr :8000 \
--model deepseek-v4-flash \
--model-map "gpt-4=deepseek-v4-pro:openai,*=deepseek-v4-flash:openai"The :openai protocol override declares the downstream speaks OpenAI Chat Completions. Responses API detection runs before the passthrough check, so the request still gets full conversion (Responses → Chat); on the response path it prevents the Chat response from being wrongly converted to Anthropic format on the way back.
2. Configure GOST to intercept Codex CLI's API calls and forward them through the converter:
# gost.yaml
services:
- name: codex-cli-proxy
addr: :8787
handler:
type: tcp
metadata:
sniffing: true
listener:
type: tcp
forwarder:
nodes:
- name: opencode-go
addr: opencode.ai:443
tls:
secure: true
serverName: opencode.ai
http:
host: opencode.ai
rewriteURL:
# Responses API /v1/responses → Chat Completions /v1/chat/completions
- match: /v1/responses
replacement: /zen/go/v1/chat/completions
requestHeader:
Authorization: "Bearer your-oc-apikey"
x-api-key: "your-oc-apikey"
rewriteRequestBody:
- rewriter: llm-converter
type: application/json
rewriteResponseBody:
- rewriter: llm-converter
type: "*"
rewriters:
- name: llm-converter
plugin:
type: http
addr: http://127.0.0.1:8000/rewrite3. Point Codex CLI at the proxy:
export OPENAI_BASE_URL=http://127.0.0.1:8787/v1
codexCodex CLI sends Responses API requests to /v1/responses; GOST intercepts them, the converter rewrites the body to Chat Completions format (with model name remapping), and the request is forwarded to opencode-go with the URL rewritten to /zen/go/v1/chat/completions. Upstream Chat Completions responses are converted back to Responses API format transparently.
This setup routes Claude Code (Anthropic protocol) directly to the Google Gemini API through GOST + llm-api-converter, with direct Anthropic↔Gemini protocol conversion:
Claude Code → GOST (proxy) → llm-api-converter → Google Gemini API
1. Start the converter with :gemini protocol override:
./llm-api-converter \
--addr :8000 \
--model-map "*=gemini-3.1-flash-lite:gemini"The :gemini protocol tells the converter the downstream speaks Gemini generateContent — since Claude Code sends Anthropic format, the converter runs Anthropic→Gemini on the request and Gemini→Anthropic on the response.
2. Configure GOST to point directly at the Gemini API:
# gost.yaml
services:
- name: claude-code-proxy
addr: :8787
handler:
type: tcp
metadata:
sniffing: true
listener:
type: tcp
forwarder:
hop: hop-0
hops:
- name: hop-0
nodes:
- name: gemini
addr: generativelanguage.googleapis.com:443
tls:
secure: true
serverName: generativelanguage.googleapis.com
http:
host: generativelanguage.googleapis.com
rewriteURL:
- match: /v1/messages
replacement: /v1beta/models/gemini-3.1-flash-lite:streamGenerateContent?alt=sse
requestHeader:
X-goog-api-key: "your-gemini-api-key"
Authorization: ""
rewriteRequestBody:
- rewriter: gemini-converter
rewriteResponseBody:
- rewriter: llm-converter
rewriters:
- name: gemini-converter
plugin:
type: http
addr: http://127.0.0.1:8000/rewriteKey details:
rewriteURLmaps Anthropic/v1/messagesto Gemini'sstreamGenerateContentwith SSE outputX-goog-api-keyis the Gemini API auth method (notAuthorization: Bearer)Authorization: ""— empty value strips the Anthropic Bearer token so it doesn't reach Google:geminiprotocol in the model-map triggers Anthropic↔Gemini body conversion in both directions
3. Point Claude Code at the proxy:
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claudeClaude Code sends Anthropic POST /v1/messages to the proxy. GOST intercepts it, the converter rewrites the body to Gemini generateContent format, and the request is forwarded directly to the Google Gemini API. Streaming responses come back as Gemini SSE chunks and are converted to Anthropic SSE events transparently.
| Direction | Description |
|---|---|
| OpenAI Request → Anthropic Request | For forwarding to Anthropic API |
| OpenAI Request → Gemini Request | For forwarding to Google Gemini API |
| OpenAI Response → Anthropic Response | For returning Anthropic-format responses to clients |
| OpenAI Response → Gemini Response | For returning Gemini-format responses to clients |
| Anthropic Request → OpenAI Request | For forwarding to OpenAI-compatible downstreams (DeepSeek, etc.) |
| Anthropic Request → Gemini Request | For forwarding to Google Gemini API from Anthropic SDK clients |
| Anthropic Response → OpenAI Response | For returning OpenAI-format responses to clients |
| Anthropic Response → Gemini Response | For returning Gemini-format responses to clients |
| Gemini Request → OpenAI Request | Reverse: Gemini format → OpenAI Chat format |
| Gemini Request → Anthropic Request | Reverse: Gemini format → Anthropic format |
| Gemini Response → OpenAI Response | Reverse: Gemini response → OpenAI Chat response |
| Gemini Response → Anthropic Response | Reverse: Gemini response → Anthropic response |
| Responses API Request → Chat Completions Request | For forwarding Codex CLI (Responses API) to OpenAI Chat Completions backends |
| Responses API Request → Anthropic Request | When the model-map routes a Responses request to an Anthropic downstream |
| Chat Completions Response → Responses API Response | Converting upstream Chat response back to Responses API format |
| Anthropic Response → Responses API Response | Converting upstream Anthropic response back to Responses API format |
Anthropic SSE — Converts OpenAI streaming delta chunks into the proper Anthropic SSE event sequence:
message_start → ping → content_block_start → content_block_delta* → content_block_stop → message_delta → message_stop
Supports text, thinking (reasoning), and tool call deltas with proper content block transitions, signature_delta for thinking blocks, and tool name restriction to prevent tool hallucination.
Gemini SSE → OpenAI — Converts Gemini streamGenerateContent SSE chunks (complete GeminiChatResponse JSON per chunk) into OpenAI delta chunks with finish_reason + [DONE] markers.
Gemini SSE → Anthropic — Converts Gemini SSE chunks into the Anthropic SSE event sequence, tracking content block transitions (text ↔ tool_use).
Responses API SSE — Converts Chat Completions streaming deltas into the Responses API SSE event sequence:
response.created → response.in_progress → output_item.added → content_part.added → response.output_text.delta* → response.output_text.done → output_item.finished → response.completed
Handles streaming reasoning content (thinking is not a first-class Responses API concept; reasoning is merged as response.output_text.delta with a type: "reasoning" annotation), text deltas, tool call accumulation across chunks, and error propagation.
Handles DeepSeek V4's requirement that reasoning_content must be preserved when tool calls are present. The cache stores reasoning across three tiers:
- Tool call ID — exact tool call replay
- Tool context — same tool pattern across different IDs
- Assistant text — text-based fallback
With optional file persistence, 30-day TTL, and FIFO eviction.
The cache backend is pluggable via the ReasoningStore interface (Get, Set, Delete, Len), allowing custom storage implementations beyond the default in-memory map.
- Pairs tool calls with their results, drops unfulfilled calls
- Converts orphan tool results to user text messages
- Merges consecutive assistant tool call messages (Claude Code conversation compression)
- Injects placeholder reasoning when DeepSeek V4 requires it but cache is empty
- Text and multi-part content blocks
- Image data URIs (
data:image/...;base64,...) - Tool use / tool result blocks
- Extended thinking / reasoning content
- System messages
- Gemini function call / function response parts
- Gemini inline data (base64 media)
- Gemini executable code / code execution result
| Flag | Default | Description |
|---|---|---|
--addr |
:8000 |
Listening address |
--model |
deepseek-chat |
Default fallback model ID |
--max-tokens |
8192 |
Default max_tokens |
--model-map |
`` | Model mapping: prefix=target[:protocol],... (* for catch-all, protocol: openai|anthropic|gemini) |
--cache |
memory |
Reasoning cache backend: memory or file:<path> |
--filter-redacted-thinking |
false |
Strip redacted_thinking blocks from Anthropic responses (OpenRouter compat) |
--log.level |
info |
Log level |
--log.format |
json |
Log format (text or json) |
llm-api-converter/
├── main.go # Entry point
├── cmd/root.go # Cobra CLI
├── convert/ # Core conversion logic
│ ├── types.go # Data types for OpenAI, Anthropic, Gemini, Responses API, SSE
│ ├── convert.go # Entry point: Convert + ConvertSSE dispatch
│ ├── detect.go # Body-primary protocol detection (positive markers)
│ ├── protocol.go # Protocol type + URI fallback + resolveModel
│ ├── registry.go # ConversionKey → converter function map
│ ├── session.go # SessionStore (per-session state, FIFO eviction)
│ ├── anthropic_to_openai.go # Anthropic → OpenAI Chat Completions
│ ├── openai_to_anthropic.go # OpenAI Chat Completions → Anthropic
│ ├── responses.go # Responses API ↔ Chat Completions
│ ├── stream.go # SSE stream utilities
│ ├── stream_anthropic.go # Anthropic SSE state machine (OpenAI → Anthropic streaming)
│ ├── stream_responses.go # Responses API SSE state machine
│ ├── stream_gemini.go # Gemini SSE → OpenAI/Anthropic streaming converters
│ ├── gemini_types.go # Gemini generateContent protocol types
│ ├── openai_to_gemini.go # OpenAI Chat ↔ Gemini (request + response)
│ ├── anthropic_to_gemini.go # Anthropic ↔ Gemini (request + response)
│ ├── reasoning_cache.go # 3-tier reasoning cache + ReasoningStore interface
│ └── gemini_test.go (and *_test.go) # Tests
├── rewriter/
│ ├── server.go # HTTP plugin server
│ └── server_test.go
├── tests/e2e/ # Integration tests
└── docs/plans/ # Historical design documents
go test ./... -v -count=1
go test ./... -race
go test ./tests/e2e/ -v -timeout 5m- deepseek-v4-opencode-claude-code-bridge — DeepSeek V4 adapter for OpenCode and Claude Code
- opencode-cc — OpenCode Claude Code bridge
- cc-switch — Claude Code provider/config switcher
- one-api — Multi-vendor LLM API management platform with model routing and key management
Part of the GOST project.