A self-hosted gateway that turns your Codex subscriptions into one OpenAI Responses endpoint with named, revocable keys.
Quick start · First release · How it works · Docs · Contributing
Important
CLAN is an independent project and is not affiliated with OpenAI. Connect only accounts you are allowed to use, and make sure the way you share access follows your provider's terms.
|
🔌 One endpoint for every client |
🔄 Several accounts act as one |
|
🔑 Access you can take back |
🌊 Faithful responses |
|
🔒 Secrets stay secret |
📦 Nothing else to run |
Usage reports: inspect generation outcomes and token usage by key, account, model or day through the admin API. Records exclude prompts and responses and are kept for 90 days by default.
- Free and open forever: every feature is open source, with no paid tiers or proprietary editions.
- Small working scope: finish the supported client scenarios before adding protocols, tools or infrastructure.
- Predictable behavior: keep the meaning of supported requests, retry only when safe and reject unsupported options explicitly.
- Clear failures: report safe error details and keep unknown usage distinct from zero.
You need Docker with Compose, just, openssl and
a Codex subscription. To run without Docker, see Running CLAN.
1. Configure and start the gateway. Put two secrets in .env and keep them
private. Restarts need the same encryption key, otherwise stored credentials cannot
be read.
git clone https://github.com/deyna256/clan && cd clan
cp .env.example .env && chmod 600 .env # then set both secrets in .env
just build && just run # serves 127.0.0.1:8080; data stays in the clan_data volume
set -a; . ./.env; set +a # load the admin token into this shell2. Connect a Codex account. Start a login, then open the returned
authorization_url in your browser.
curl http://127.0.0.1:8080/api/oauth/login \
-H "Authorization: Bearer $CLAN_ADMIN_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"name":"Primary"}'3. Create a client key. Save the returned key: CLAN shows it only once.
curl http://127.0.0.1:8080/api/client-keys \
-H "Authorization: Bearer $CLAN_ADMIN_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"name":"My laptop","concurrency_limit":3}'4. Generate. Use any model ID from GET /v1/models.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="<CLAN key>", max_retries=0)
reply = client.responses.create(model="MODEL_ID", input="Say hello.", store=False)
print(reply.status, reply.output_text)Check status before using the output: HTTP 200 also covers incomplete and
failed responses. Running CLAN covers configuration, remote
sign-in and shutdown. The client contract has
the OpenCode configuration.
Everything below is implemented. Live checks with OpenCode 1.18.29 and the OpenAI Python SDK 3.13.0, known limitations and what is not checked yet are listed in the client contract.
| Area | Scope | Status |
|---|---|---|
| Clients | OpenCode and Python applications using the OpenAI SDK | ✅ |
| Client API | POST /v1/responses for generation; GET /v1/models for available models |
✅ |
| Integration | Codex subscription access through built-in OAuth sign-in | ✅ |
| Generation | Text, instructions, client-supplied history, image input including screenshots, JSON and JSON Schema output, reasoning settings and continuation data | ✅ |
| Tools | Function definitions, calls, JSON argument fragments and results; tools run in the client | ✅ |
| Responses | Complete JSON responses and SSE streaming | ✅ |
| Accounts | One or more Codex accounts; round-robin for the requested model; temporarily exclude accounts with exhausted provider limits | ✅ |
| Client access | Separate named keys for applications or people, with shared access to connected accounts and available models | ✅ |
| Limits | Concurrent requests per key, with immediate rejection when full | ✅ |
| Management | HTTP JSON API under /api, protected by a separate admin token |
✅ |
| Storage | SQLite for accounts, encrypted OAuth credentials, access-key hashes and settings | ✅ |
| Diagnostics | Structured JSON logs with outcomes, safe errors, account switches and known token usage | ✅ |
| Deployment | One Go process or container per installation | ✅ |
Clients send the full conversation history with each request. CLAN does not store
responses or support continuation through previous_response_id. For Codex
compatibility, CLAN removes max_output_tokens and logs a warning, so the requested
output limit is not enforced.
Revoking a client key blocks new requests and cancels its active requests. Retries keep the same concurrency slot, and resources are closed before the slot is released. Known token usage is diagnostic data, not a token budget. The architecture defines management operations, OAuth setup and remaining integration checks.
Outside the first release
- Chat Completions, Anthropic and Gemini APIs; other upstream integrations and API-key authentication to providers.
- Provider-native tools (web search, code execution, image generation and hosted MCP), custom/freeform tools, namespaces and tool search.
- WebSocket, background generation, stored responses, Conversations, Files, Containers, Vector Stores and separate compact/count operations.
- Per-key account/model permissions, RPM limits, token or monetary budgets, database request history and usage reports.
- PostgreSQL, multiple active gateway instances, a web panel, auth-file imports and additional OAuth connection methods.
These exclusions do not commit the project to a later delivery date.
This diagram shows runtime flow inside one process, not Go package dependencies. Solid arrows carry requests; dashed arrows carry responses.
flowchart TB
accTitle: CLAN first-release request flow
accDescr: OpenCode or a Python application sends a Responses request to CLAN. The HTTP API validates it, request execution manages access and attempts, and the Codex integration calls Codex using OAuth. Responses return through the same modules.
client["OpenCode / Python OpenAI SDK"]
subgraph clan["CLAN · one process"]
http["HTTP API<br/>Validate Responses input<br/>Write JSON or SSE"]
execution["Request execution<br/>Check key and concurrency<br/>Select account and manage attempts"]
codex["Codex integration<br/>Authenticate with OAuth<br/>Adapt protocol and classify outcomes"]
end
upstream["Codex"]
client -->|"Responses + CLAN key"| http
http -->|"Validated request + CLAN key"| execution
execution -->|"Request + selected account"| codex
codex -->|"Authenticated request"| upstream
upstream -.->|"Provider response / stream"| codex
codex -.->|"Responses data + execution metadata"| execution
execution -.->|"Response / stream / error"| http
http -.->|"Responses JSON / SSE / error"| client
Responses is the reference content format. Execution uses small metadata values without parsing messages or tool arguments. The Codex integration handles protocol differences; there is no separate universal message or event model.
| I want to… | Read |
|---|---|
| Start the gateway and connect an account | Running CLAN |
| Connect OpenCode or the Python SDK and check compatibility | Client contract |
| Manage accounts, client keys and model lists | Management API |
| Understand modules, management and open questions | Architecture |
| Write and review Go code and tests | Development guide |
| Open an issue or PR, or update documentation | Contributing |
| Check community rules | Code of Conduct |
Decision log
These are accepted design decisions. Their status does not imply implementation.
| Decision | Covers |
|---|---|
| Use Go | Application language |
| Use Responses as the content format | Responses content and small execution metadata |
| Separate execution from protocols | Attempt ownership, concurrency, cancellation and streaming |
| Use SQLite | Persistent configuration and secret storage |
| Use chi | HTTP routing over net/http |
Built with Go and net/http, chi for routing,
Huma for the management OpenAPI contract and
modernc SQLite.
Issues and pull requests are welcome. Read CONTRIBUTING.md for
the issue format, branch and commit rules, and the checks to run. Working notes
belong in the ignored .local/ directory; see
local working documents.
MIT © 2026 Ivan Deyna