Skip to content

Repository files navigation

llm-api-converter

English | 简体中文

A GOST Rewriter HTTP plugin that converts bidirectionally between OpenAI Chat Completions, Anthropic Messages, and Google Gemini generateContent API formats. Designed for use with tools like Claude Code, Codex CLI, OpenCode, and other LLM clients that speak any of these protocols.

Table of Contents

How it works

Deployed as a GOST rewriter plugin, it intercepts HTTP request/response bodies in a forward proxy and transparently converts between the two wire formats:

Anthropic SDK client → GOST → llm-api-converter → OpenAI-compatible API
                    ↕ (protocol conversion)       ↕
                Anthropic format               OpenAI format

The converter auto-detects the input format using positive structural markers only — no negative exclusions, no hardcoded model prefixes. A SessionStore tracks client protocol across request/response pairs for correct bidirectional routing. A URI-based fallback handles minimal requests that lack distinguishing features.

Quick start

# Build
go build -o llm-api-converter .

# Run standalone
./llm-api-converter --addr :8000 --model-map "claude-opus=deepseek-v4-pro:openai,claude-sonnet=deepseek-v4-flash,*=deepseek-v4-flash:openai"

# Run with GOST
gost -C gost.yaml

With Docker Compose

Run the converter as a container alongside GOST. The published image is ginuerzh/llm-api-converter (multi-arch: amd64/arm64/arm v6/v7). Since the image ENTRYPOINT is the binary, command: supplies the CLI flags.

# docker-compose.yml
services:
  llm-converter:
    image: ginuerzh/llm-api-converter:latest
    command:
      - --addr
      - :8000
      - --model
      - deepseek-v4-flash
      - --model-map
      - claude-opus=deepseek-v4-pro:openai,claude-sonnet=deepseek-v4-flash,*=deepseek-v4-flash:openai
    ports:
      - "8000:8000"
    restart: unless-stopped

  gost:
    image: gogost/gost:latest
    command: -C /etc/gost/gost.yaml
    volumes:
      - ./gost.yaml:/etc/gost/gost.yaml:ro
    ports:
      - "8787:8787"
    depends_on:
      - llm-converter
    restart: unless-stopped

Point the GOST rewriter plugin at the converter's container address:

# in gost.yaml
rewriters:
- name: llm-converter
  plugin:
    type: http
    addr: http://llm-converter:8000/rewrite
docker compose up -d
export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claude

Build the image locally instead of pulling:

docker build -t ginuerzh/llm-api-converter .
# or, with the multi-arch buildx workflow from .github/workflows/buildx.yml

Claude Code → DeepSeek (via opencode-go)

This setup lets Claude Code (Anthropic protocol) call DeepSeek models (OpenAI protocol) through the converter:

Claude Code → GOST (proxy) → llm-api-converter → opencode-go API → DeepSeek

1. Start the converter:

./llm-api-converter \
  --addr :8000 \
  --model deepseek-v4-flash \
  --model-map "claude-opus=deepseek-v4-pro:openai,*=deepseek-v4-flash:openai"

2. Configure GOST to intercept Anthropic API calls and forward them through the converter:

# gost.yaml
services:
- name: claude-code-proxy
  addr: :8787
  handler:
    type: tcp
    metadata:
      sniffing: true
  listener:
    type: tcp
  forwarder:
    nodes:
    - name: opencode-go
      addr: opencode.ai:443
      tls:
        secure: true
        serverName: opencode.ai
      http:
        host: opencode.ai
        rewriteURL:
        # Anthropic /v1/messages → OpenAI /v1/chat/completions
        - match: /v1/messages
          replacement: /zen/go/v1/chat/completions
        requestHeader:
          Authorization: "Bearer your-oc-apikey"
          x-api-key: "your-oc-apikey"
        rewriteRequestBody:
        - rewriter: llm-converter
          type: application/json
        rewriteResponseBody:
        - rewriter: llm-converter
          type: "*"

rewriters:
- name: llm-converter
  plugin:
    type: http
    addr: http://127.0.0.1:8000/rewrite

3. Point Claude Code at the proxy:

export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claude

All Anthropic traffic from Claude Code is intercepted by GOST, converted to OpenAI Chat Completions format by the plugin, and forwarded to the opencode-go API for DeepSeek inference. Responses and SSE streams are converted back to Anthropic format transparently.

Model mapping notes:

  • claude-opus=deepseek-v4-pro:openai: Routes requests with model name starting with claude-opus to DeepSeek V4 Pro, converting Anthropic→OpenAI
  • *=deepseek-v4-flash:openai: Catch-all fallback for any unmatched model prefix, also converting to OpenAI
  • Downstream protocol override: Append :openai, :anthropic, or :gemini after the target to declare what format the backend speaks (prefix=target:protocol). Without a protocol suffix, the default is passthrough — the body passes through with only the model name rewritten, no format conversion. With :openai/:anthropic/:gemini, conversion runs when the incoming protocol differs from the declared one; when they match, only the model is rewritten. The override applies on both request and response paths via per-session client protocol tracking. Example: claude-opus=deepseek-v4-pro:openai — incoming Anthropic differs from :openai, so Anthropic→OpenAI conversion runs; claude-opus=deepseek-v4-pro:anthropic — incoming Anthropic matches, so only the model is rewritten.
  • Note: :responses is not a valid override (only openai/anthropic/gemini); Responses API traffic is detected and routed via body markers and the session store, not the model map. Empty targets (e.g. claude-opus=:openai) are rejected at parse time.

Update the --model-map to match your opencode-go deployment's available models.

LLM Router (body routing)

A complete multi-vendor LLM router setup using GOST + llm-api-converter — routes requests to different providers based on prompt content, with automatic protocol conversion. See the LLM Router blog post for the full guide.

Codex CLI → DeepSeek (via opencode-go)

This setup lets Codex CLI (OpenAI Responses API protocol) call DeepSeek models (OpenAI Chat Completions protocol) through the converter:

Codex CLI → GOST (proxy) → llm-api-converter → opencode-go API → DeepSeek

Codex CLI sends Responses API format (POST /v1/responses with {model, input, ...}); the converter translates to Chat Completions format (POST /v1/chat/completions with {model, messages, ...}) for opencode-go, and reverses the response on the way back.

1. Start the converter:

./llm-api-converter \
  --addr :8000 \
  --model deepseek-v4-flash \
  --model-map "gpt-4=deepseek-v4-pro:openai,*=deepseek-v4-flash:openai"

The :openai protocol override declares the downstream speaks OpenAI Chat Completions. Responses API detection runs before the passthrough check, so the request still gets full conversion (Responses → Chat); on the response path it prevents the Chat response from being wrongly converted to Anthropic format on the way back.

2. Configure GOST to intercept Codex CLI's API calls and forward them through the converter:

# gost.yaml
services:
- name: codex-cli-proxy
  addr: :8787
  handler:
    type: tcp
    metadata:
      sniffing: true
  listener:
    type: tcp
  forwarder:
    nodes:
    - name: opencode-go
      addr: opencode.ai:443
      tls:
        secure: true
        serverName: opencode.ai
      http:
        host: opencode.ai
        rewriteURL:
        # Responses API /v1/responses → Chat Completions /v1/chat/completions
        - match: /v1/responses
          replacement: /zen/go/v1/chat/completions
        requestHeader:
          Authorization: "Bearer your-oc-apikey"
          x-api-key: "your-oc-apikey"
        rewriteRequestBody:
        - rewriter: llm-converter
          type: application/json
        rewriteResponseBody:
        - rewriter: llm-converter
          type: "*"

rewriters:
- name: llm-converter
  plugin:
    type: http
    addr: http://127.0.0.1:8000/rewrite

3. Point Codex CLI at the proxy:

export OPENAI_BASE_URL=http://127.0.0.1:8787/v1
codex

Codex CLI sends Responses API requests to /v1/responses; GOST intercepts them, the converter rewrites the body to Chat Completions format (with model name remapping), and the request is forwarded to opencode-go with the URL rewritten to /zen/go/v1/chat/completions. Upstream Chat Completions responses are converted back to Responses API format transparently.

Claude Code → Gemini

This setup routes Claude Code (Anthropic protocol) directly to the Google Gemini API through GOST + llm-api-converter, with direct Anthropic↔Gemini protocol conversion:

Claude Code → GOST (proxy) → llm-api-converter → Google Gemini API

1. Start the converter with :gemini protocol override:

./llm-api-converter \
  --addr :8000 \
  --model-map "*=gemini-3.1-flash-lite:gemini"

The :gemini protocol tells the converter the downstream speaks Gemini generateContent — since Claude Code sends Anthropic format, the converter runs Anthropic→Gemini on the request and Gemini→Anthropic on the response.

2. Configure GOST to point directly at the Gemini API:

# gost.yaml
services:
- name: claude-code-proxy
  addr: :8787
  handler:
    type: tcp
    metadata:
      sniffing: true
  listener:
    type: tcp
  forwarder:
    hop: hop-0

hops:
- name: hop-0
  nodes:
  - name: gemini
    addr: generativelanguage.googleapis.com:443
    tls:
      secure: true
      serverName: generativelanguage.googleapis.com
    http:
      host: generativelanguage.googleapis.com
      rewriteURL:
      - match: /v1/messages
        replacement: /v1beta/models/gemini-3.1-flash-lite:streamGenerateContent?alt=sse
      requestHeader:
        X-goog-api-key: "your-gemini-api-key"
        Authorization: ""
      rewriteRequestBody:
      - rewriter: gemini-converter
      rewriteResponseBody:
      - rewriter: llm-converter

rewriters:
- name: gemini-converter
  plugin:
    type: http
    addr: http://127.0.0.1:8000/rewrite

Key details:

  • rewriteURL maps Anthropic /v1/messages to Gemini's streamGenerateContent with SSE output
  • X-goog-api-key is the Gemini API auth method (not Authorization: Bearer)
  • Authorization: "" — empty value strips the Anthropic Bearer token so it doesn't reach Google
  • :gemini protocol in the model-map triggers Anthropic↔Gemini body conversion in both directions

3. Point Claude Code at the proxy:

export ANTHROPIC_BASE_URL=http://127.0.0.1:8787
claude

Claude Code sends Anthropic POST /v1/messages to the proxy. GOST intercepts it, the converter rewrites the body to Gemini generateContent format, and the request is forwarded directly to the Google Gemini API. Streaming responses come back as Gemini SSE chunks and are converted to Anthropic SSE events transparently.

Capabilities

Protocol conversion

Direction Description
OpenAI Request → Anthropic Request For forwarding to Anthropic API
OpenAI Request → Gemini Request For forwarding to Google Gemini API
OpenAI Response → Anthropic Response For returning Anthropic-format responses to clients
OpenAI Response → Gemini Response For returning Gemini-format responses to clients
Anthropic Request → OpenAI Request For forwarding to OpenAI-compatible downstreams (DeepSeek, etc.)
Anthropic Request → Gemini Request For forwarding to Google Gemini API from Anthropic SDK clients
Anthropic Response → OpenAI Response For returning OpenAI-format responses to clients
Anthropic Response → Gemini Response For returning Gemini-format responses to clients
Gemini Request → OpenAI Request Reverse: Gemini format → OpenAI Chat format
Gemini Request → Anthropic Request Reverse: Gemini format → Anthropic format
Gemini Response → OpenAI Response Reverse: Gemini response → OpenAI Chat response
Gemini Response → Anthropic Response Reverse: Gemini response → Anthropic response
Responses API Request → Chat Completions Request For forwarding Codex CLI (Responses API) to OpenAI Chat Completions backends
Responses API Request → Anthropic Request When the model-map routes a Responses request to an Anthropic downstream
Chat Completions Response → Responses API Response Converting upstream Chat response back to Responses API format
Anthropic Response → Responses API Response Converting upstream Anthropic response back to Responses API format

Streaming

Anthropic SSE — Converts OpenAI streaming delta chunks into the proper Anthropic SSE event sequence:

message_start → ping → content_block_start → content_block_delta* → content_block_stop → message_delta → message_stop

Supports text, thinking (reasoning), and tool call deltas with proper content block transitions, signature_delta for thinking blocks, and tool name restriction to prevent tool hallucination.

Gemini SSE → OpenAI — Converts Gemini streamGenerateContent SSE chunks (complete GeminiChatResponse JSON per chunk) into OpenAI delta chunks with finish_reason + [DONE] markers.

Gemini SSE → Anthropic — Converts Gemini SSE chunks into the Anthropic SSE event sequence, tracking content block transitions (text ↔ tool_use).

Responses API SSE — Converts Chat Completions streaming deltas into the Responses API SSE event sequence:

response.created → response.in_progress → output_item.added → content_part.added → response.output_text.delta* → response.output_text.done → output_item.finished → response.completed

Handles streaming reasoning content (thinking is not a first-class Responses API concept; reasoning is merged as response.output_text.delta with a type: "reasoning" annotation), text deltas, tool call accumulation across chunks, and error propagation.

Multi-tier reasoning cache (DeepSeek V4)

Handles DeepSeek V4's requirement that reasoning_content must be preserved when tool calls are present. The cache stores reasoning across three tiers:

  1. Tool call ID — exact tool call replay
  2. Tool context — same tool pattern across different IDs
  3. Assistant text — text-based fallback

With optional file persistence, 30-day TTL, and FIFO eviction.

The cache backend is pluggable via the ReasoningStore interface (Get, Set, Delete, Len), allowing custom storage implementations beyond the default in-memory map.

Message sequence sanitization

  • Pairs tool calls with their results, drops unfulfilled calls
  • Converts orphan tool results to user text messages
  • Merges consecutive assistant tool call messages (Claude Code conversation compression)
  • Injects placeholder reasoning when DeepSeek V4 requires it but cache is empty

Content support

  • Text and multi-part content blocks
  • Image data URIs (data:image/...;base64,...)
  • Tool use / tool result blocks
  • Extended thinking / reasoning content
  • System messages
  • Gemini function call / function response parts
  • Gemini inline data (base64 media)
  • Gemini executable code / code execution result

CLI flags

Flag Default Description
--addr :8000 Listening address
--model deepseek-chat Default fallback model ID
--max-tokens 8192 Default max_tokens
--model-map `` Model mapping: prefix=target[:protocol],... (* for catch-all, protocol: openai|anthropic|gemini)
--cache memory Reasoning cache backend: memory or file:<path>
--filter-redacted-thinking false Strip redacted_thinking blocks from Anthropic responses (OpenRouter compat)
--log.level info Log level
--log.format json Log format (text or json)

Project structure

llm-api-converter/
├── main.go              # Entry point
├── cmd/root.go          # Cobra CLI
├── convert/             # Core conversion logic
│   ├── types.go                              # Data types for OpenAI, Anthropic, Gemini, Responses API, SSE
│   ├── convert.go                            # Entry point: Convert + ConvertSSE dispatch
│   ├── detect.go                             # Body-primary protocol detection (positive markers)
│   ├── protocol.go                           # Protocol type + URI fallback + resolveModel
│   ├── registry.go                           # ConversionKey → converter function map
│   ├── session.go                            # SessionStore (per-session state, FIFO eviction)
│   ├── anthropic_to_openai.go                # Anthropic → OpenAI Chat Completions
│   ├── openai_to_anthropic.go                # OpenAI Chat Completions → Anthropic
│   ├── responses.go                          # Responses API ↔ Chat Completions
│   ├── stream.go                             # SSE stream utilities
│   ├── stream_anthropic.go                   # Anthropic SSE state machine (OpenAI → Anthropic streaming)
│   ├── stream_responses.go                   # Responses API SSE state machine
│   ├── stream_gemini.go                      # Gemini SSE → OpenAI/Anthropic streaming converters
│   ├── gemini_types.go                       # Gemini generateContent protocol types
│   ├── openai_to_gemini.go                   # OpenAI Chat ↔ Gemini (request + response)
│   ├── anthropic_to_gemini.go                # Anthropic ↔ Gemini (request + response)
│   ├── reasoning_cache.go                    # 3-tier reasoning cache + ReasoningStore interface
│   └── gemini_test.go (and *_test.go)        # Tests
├── rewriter/
│   ├── server.go                             # HTTP plugin server
│   └── server_test.go
├── tests/e2e/                                # Integration tests
└── docs/plans/                               # Historical design documents

Tests

go test ./... -v -count=1
go test ./... -race
go test ./tests/e2e/ -v -timeout 5m

Related projects

License

Part of the GOST project.

About

A GOST Rewriter HTTP plugin that converts bidirectionally between OpenAI Chat Completions and Anthropic Messages API formats

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages