Turn your Claude Max / OpenAI Pro subscriptions into a private API endpoint that any project can consume.
- Problem Statement
- Why CLI Wrappers Instead of SDKs
- Architecture Overview
- CLI Capability Matrix
- Project Structure
- Security Measures
- Key Management
- Step-by-Step Deployment
- Using the Bridge From Other Projects
- API Reference
- Maintenance & Operations
AI API calls are expensive at scale. If you already pay for Claude Max ($100-200/mo) or OpenAI Pro ($200/mo), you're paying for generous usage that's locked to the CLI tools (Claude Code CLI, Codex CLI). These CLIs can only run locally from a terminal — they can't be called from web apps, Figma plugins, mobile apps, or any HTTP-based client.
This project bridges that gap: it wraps the CLIs behind an HTTP API, deploys to a cheap server, and exposes a URL that any project can use as if it were a regular AI API — backed by your subscription instead of per-token billing.
The obvious question: why not use the Anthropic SDK (@anthropic-ai/sdk) or OpenAI SDK directly?
Because SDKs require API keys billed per-token. The entire point is to use your existing subscription. The CLI tools authenticate via OAuth to your Max/Pro account, and usage counts against your subscription's monthly allocation — not a separate API bill.
| Approach | Auth | Billing | Setup |
|---|---|---|---|
| SDK (Anthropic/OpenAI) | API key | Per-token ($$$) | Simple |
| CLI wrapper (this project) | OAuth subscription | Monthly flat rate | More setup, but free per-request |
The tradeoff: more infrastructure complexity in exchange for dramatically lower marginal cost per request.
┌─────────────────────────────────────────────────────────────────┐
│ Your Other Projects (web apps, plugins, scripts, etc.) │
│ fetch('https://bridge.yourdomain.com/generate', { │
│ headers: { Authorization: 'Bearer <USER_KEY>' }, │
│ body: { systemPrompt, userPrompt, model } │
│ }) │
└────────────────────────────┬────────────────────────────────────┘
│ HTTPS
▼
┌─────────────────────────────────────────────────────────────────┐
│ Cloudflare Tunnel │
│ - Free TLS termination │
│ - DDoS protection │
│ - No open ports on server │
│ - Custom domain (bridge.yourdomain.com) │
└────────────────────────────┬────────────────────────────────────┘
│ HTTP (localhost only)
▼
┌─────────────────────────────────────────────────────────────────┐
│ VPS (e.g. DigitalOcean Droplet, $4-6/mo) │
│ ┌───────────────────────────────────────────────────────────┐ │
│ │ ai-cli-bridge (Express server, managed by systemd) │ │
│ │ - Per-user key auth (SHA-256 hashed, timing-safe) │ │
│ │ - Per-key usage limits (requests, tokens, cost) │ │
│ │ - Rate limiting, security headers, input validation │ │
│ │ - Admin API for key management │ │
│ └──────────┬──────────────────────────┬─────────────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌──────────────────┐ ┌──────────────────────┐ │
│ │ Claude Code CLI │ │ Codex CLI │ │
│ │ (OAuth → Max) │ │ (OAuth → Pro) │ │
│ └──────────────────┘ └──────────────────────┘ │
└─────────────────────────────────────────────────────────────────┘
- Any cheap VPS ($4-6/mo): The CLIs call remote APIs, so CPU/RAM needs are minimal. DigitalOcean, Hetzner, Linode, etc. all work.
- Cloudflare Tunnel: Eliminates the need to open any ports on the droplet. All traffic flows through Cloudflare's network. Free TLS, free DDoS protection, and a clean domain name.
- systemd: Native Linux service manager — auto-restarts the server on crash, starts on boot, and provides logging via
journalctl. - Express: Minimal HTTP framework. The server is ~260 lines — just middleware, validation, and CLI invocation.
While a Dockerfile is included for reproducibility, the recommended deployment is directly on the droplet (not containerized). Reason: the CLIs require interactive OAuth authentication, which is easier to do via SSH than inside a container. Once authenticated, the tokens persist on the filesystem.
Important to understand what you get (and don't get) using CLI vs API:
| Capability | Claude Code CLI (Max) | Claude API (paid) | Codex CLI (Pro) |
|---|---|---|---|
| Extended thinking | Automatic | Fine-grained budget control | Automatic |
| Web search | No native tool | web_search tool |
No native tool |
| Vision/images | Supported | Supported | Supported |
| Tool use | Bash, file read/write | Custom tool definitions | Sandbox-based |
| Streaming | Possible (via spawn()) |
Native SSE | Possible |
| Model selection | Limited to subscription tier | Any model | Limited to subscription |
| Usage limits | Monthly subscription cap | Pay-per-token, no cap | Monthly cap |
| Prompt caching | Automatic | Manual API params | Automatic |
Key limitations: No web search, no thinking budget control, subscription usage caps apply.
ai-cli-bridge/
├── src/
│ ├── server.ts # Express app — security headers, CORS, rate
│ │ # limiting, admin routes, generation routes
│ ├── config.ts # All settings loaded from environment variables
│ ├── keys.ts # KeyManager — SHA-256 hashed key storage, CRUD,
│ │ # per-key usage tracking, limit enforcement
│ ├── middleware/
│ │ └── auth.ts # Per-user key auth + admin auth (timing-safe)
│ └── providers/
│ ├── claude.ts # Claude Code CLI wrapper (execFile → promise)
│ └── codex.ts # Codex CLI wrapper (JSONL parsing)
├── data/ # Runtime data (gitignored)
│ ├── keys.json # SHA-256 hashes → key config + limits
│ └── usage.json # Per-key daily/monthly usage counters
├── Dockerfile # Container build (non-root user, dumb-init)
├── docker-compose.yml # With health check + auth volume persistence
├── ai-cli-bridge.service # systemd service file
├── cloudflared-config.yml # Cloudflare Tunnel config template
├── setup.sh # One-shot droplet provisioning script
├── .env.example # All configurable environment variables
├── .gitignore
├── package.json
└── tsconfig.json
Both providers follow the same pattern:
- Receive a
{ systemPrompt, userPrompt, model }request - Write the system prompt to a temp file (with
0o600permissions — owner-only) - Spawn the CLI via
execFile()(notexec()— avoids shell injection) - Pipe the user prompt to stdin
- Parse the CLI's JSON/JSONL output
- Clean up the temp file
- Return a normalized response matching the Anthropic Messages API shape
The key insight: execFile() passes arguments as an array, not a shell string, which prevents command injection even if prompt content contains shell metacharacters.
Claude CLI invocation:
claude -p \
--system-prompt-file /tmp/cli-bridge-<uuid>.txt \
--model <model> \
--max-turns 5 \
--tools '' \
--output-format json
< userPrompt (via stdin)
Codex CLI invocation:
codex exec \
--model <model> \
--full-auto \
--sandbox read-only \
--json \
- < combinedPrompt (via stdin)
Problem: Storing raw API keys on disk means a file leak exposes all keys.
Solution: Only SHA-256 hashes are stored in data/keys.json. Raw keys are shown once at creation time and never stored. Validation hashes the incoming key and compares against stored hashes using timing-safe comparison.
// On key creation — raw key returned to admin, only hash stored
const rawKey = randomBytes(32).toString('hex');
const keyHash = createHash('sha256').update(rawKey).digest('hex');
this.keys.keys[keyHash] = { name, limits, ... };
// On validation — combined validate + limit check in single call
validateAndCheck(rawKey: string): { name?: string; hash?: string; error?: string }Problem: Standard string comparison (===) short-circuits on the first mismatched character. An attacker can measure response times to guess valid keys character-by-character.
Solution: crypto.timingSafeEqual() compares all bytes regardless of where differences occur. Used for both user key and admin key validation.
Problem: Attaching the raw API key to the Express req object (req.rawKey) means it persists in memory and could leak through error handlers, logging, or middleware.
Solution: The validateAndCheck() method returns the SHA-256 hash, which is stored as req.keyHash. The raw key exists only as a local variable during auth and is discarded immediately.
Problem: If keys.json or usage.json become corrupted (power loss during write, disk error), JSON.parse() throws and the server crashes on startup.
Solution: Both files are loaded inside try-catch blocks. On parse failure, the server starts with empty data and logs a warning rather than crashing.
Problem: Without cleanup, usage.json grows indefinitely as daily/monthly entries accumulate.
Solution: On startup, entries older than 90 days (daily) or 12 months (monthly) are automatically pruned.
Problem: writeFileSync() defaults to mode 0o666 (world-readable). System prompts written to temp files could be read by other users on the system.
Solution: Write with mode: 0o600 (owner read/write only).
Problem: Returning raw CLI error messages to clients leaks system paths, command arguments, and internal state.
Solution: Log detailed errors server-side, return generic messages to clients.
// Server-side: full details for debugging
console.error('[claude] CLI execution failed');
// Client-side: generic message
catch { res.status(500).json({ error: 'Generation failed' }); }Standard hardening headers applied to all responses:
res.setHeader('X-Content-Type-Options', 'nosniff');
res.setHeader('X-Frame-Options', 'DENY');
res.setHeader('Referrer-Policy', 'strict-origin-when-cross-origin');
res.setHeader('Permissions-Policy', 'geolocation=(), microphone=(), camera=()');All generation endpoints validate:
systemPromptanduserPromptmust be strings (prevents type confusion)- Maximum length enforced (500K chars — prevents memory exhaustion)
modelparameter must be a string if provided
express-rate-limit with configurable window and max requests. trust proxy is set to 1 so rate limiting correctly identifies clients behind Cloudflare Tunnel (which sends X-Forwarded-For).
Cloudflare Tunnel means the droplet has zero open ports. All traffic flows through Cloudflare's encrypted tunnel. Even if someone discovers the droplet's IP, there's nothing to connect to.
The Dockerfile runs as a non-root bridge user with dumb-init for proper signal handling. Auth volumes are mapped to the non-root home directory.
- HTTPS on the server itself: Not needed — Cloudflare Tunnel handles TLS termination. The server listens on HTTP internally, which is standard for reverse-proxy architectures.
- Model whitelisting: Left flexible so new models work without code changes. The CLIs themselves enforce model access based on your subscription.
- Request logging to a database: Overkill for a personal bridge. Console logs captured by journalctl are sufficient. Per-key usage is tracked in
usage.json.
The bridge supports multi-user access with per-key usage limits. An admin key (set via BRIDGE_ADMIN_KEY env var) controls key CRUD operations.
- Admin creates a user key via
POST /admin/keyswith a name and optional limits - The raw key is returned once — the admin gives it to the user
- Only the SHA-256 hash is stored on disk (
data/keys.json) - Each request validates the key, checks limits, and tracks usage
- Usage is tracked per-key with daily and monthly granularity
All limits default to 0 (unlimited). You can mix and match any combination:
| Limit | Field | Description |
|---|---|---|
| Requests per day | maxRequestsPerDay |
Hard cap on daily request count |
| Requests per month | maxRequestsPerMonth |
Hard cap on monthly request count |
| Tokens per month | maxTokensPerMonth |
Combined input + output tokens per month |
| Cost per day | maxCostPerDay |
USD spend cap per day |
| Cost per month | maxCostPerMonth |
USD spend cap per month |
Each request records:
- Request count — incremented per call
- Token count — input + output tokens combined
- Cost (USD) — actual
cost_usdfrom the provider response
Usage is tracked in-memory and flushed to disk every 30 seconds. Old entries are pruned automatically (daily > 90 days, monthly > 12 months).
If no keys exist in data/keys.json, auth is completely bypassed. This is useful for local development. Create your first key via the admin API to enable auth.
Create a key with a $5/month cost cap:
curl -X POST https://bridge.yourdomain.com/admin/keys \
-H "Authorization: Bearer <ADMIN_KEY>" \
-H "Content-Type: application/json" \
-d '{"name":"alice","maxCostPerMonth":5.00}'
# → {"key":"abc123...","note":"Save this key — it cannot be retrieved again."}Create a key with request + cost limits:
curl -X POST https://bridge.yourdomain.com/admin/keys \
-H "Authorization: Bearer <ADMIN_KEY>" \
-H "Content-Type: application/json" \
-d '{"name":"bob","maxRequestsPerDay":50,"maxCostPerMonth":10.00}'Check a user's usage:
curl https://bridge.yourdomain.com/admin/keys/alice \
-H "Authorization: Bearer <ADMIN_KEY>"
# → {"name":"alice","limits":{...},"usage":{"today":{...},"thisMonth":{...}}}Update limits on an existing key:
curl -X PATCH https://bridge.yourdomain.com/admin/keys/alice \
-H "Authorization: Bearer <ADMIN_KEY>" \
-H "Content-Type: application/json" \
-d '{"maxCostPerDay":1.00}'- A DigitalOcean account (or any VPS provider)
- A Cloudflare account with a domain
- Claude Max and/or OpenAI Pro subscription
- SSH key pair on your local machine
- Log into DigitalOcean → Create → Droplets
- Image: Ubuntu 24.04 LTS
- Plan: Basic, $4/mo (512MB RAM) or $6/mo (1GB)
- Region: Closest to you
- Auth: Add your SSH public key
- Hostname:
ai-cli-bridge - Create
Add to your ~/.ssh/config:
Host ai-bridge
HostName <DROPLET_IP>
User root
IdentityFile ~/.ssh/<YOUR_KEY>
AddKeysToAgent yes
Verify: ssh ai-bridge "echo connected"
SSH in and run:
ssh ai-bridge
# System packages
apt-get update -qq && apt-get install -y -qq curl git unzip
# Bun
curl -fsSL https://bun.sh/install | bash
source ~/.bashrc
# AI CLIs
bun install -g @anthropic-ai/claude-code @openai/codex
# Cloudflared
curl -fsSL https://github.com/cloudflare/cloudflared/releases/latest/download/cloudflared-linux-amd64.deb -o /tmp/cloudflared.deb
dpkg -i /tmp/cloudflared.deb && rm /tmp/cloudflared.debFrom your local machine:
rsync -avz --exclude node_modules --exclude .env --exclude data \
./ai-cli-bridge/ ai-bridge:/opt/ai-cli-bridge/On the droplet:
cd /opt/ai-cli-bridge
bun install
bun run build# Generate a secure admin key
ADMIN_KEY=$(openssl rand -hex 32)
echo "Your admin key: $ADMIN_KEY"
# Create .env
cat > /opt/ai-cli-bridge/.env << EOF
PORT=3456
BRIDGE_ADMIN_KEY=$ADMIN_KEY
DATA_DIR=/opt/ai-cli-bridge/data
RATE_LIMIT_WINDOW_MS=60000
RATE_LIMIT_MAX_REQUESTS=30
CORS_ORIGINS=
CLAUDE_DEFAULT_MODEL=claude-sonnet-4-20250514
CODEX_DEFAULT_MODEL=gpt-5.3-codex
EOF
chmod 600 /opt/ai-cli-bridge/.envSave the admin key — this is used to manage user keys.
This is interactive and requires a browser:
# Claude Code — follow the OAuth URL it prints
claude
# Codex — follow the OAuth URL it prints
codex authEach CLI will print a URL. Open it in your browser, authenticate, and the tokens are saved to ~/.claude/ and ~/.config/ respectively.
# Copy the service file and enable it
cp /opt/ai-cli-bridge/ai-cli-bridge.service /etc/systemd/system/
systemctl daemon-reload
systemctl enable --now ai-cli-bridgeVerify: curl -s http://localhost:3456/health → {"status":"ok"}
# Authenticate with Cloudflare (opens browser URL)
cloudflared tunnel login
# Create a named tunnel
cloudflared tunnel create ai-bridge
# Route your subdomain to the tunnel
cloudflared tunnel route dns ai-bridge bridge.yourdomain.com
# Write tunnel config
TUNNEL_ID=$(cloudflared tunnel list -o json | python3 -c "import sys,json; print(json.load(sys.stdin)[0]['id'])")
cat > ~/.cloudflared/config.yml << EOF
tunnel: $TUNNEL_ID
credentials-file: /root/.cloudflared/$TUNNEL_ID.json
ingress:
- hostname: bridge.yourdomain.com
service: http://localhost:3456
- service: http_status:404
EOF
# Install as a system service (auto-starts on reboot)
cloudflared service installcurl -X POST https://bridge.yourdomain.com/admin/keys \
-H "Authorization: Bearer <ADMIN_KEY>" \
-H "Content-Type: application/json" \
-d '{"name":"yourname","maxCostPerMonth":20.00}'Save the returned key — give it to users or use it in your projects.
From your local machine:
# Health check (unauthenticated)
curl -s https://bridge.yourdomain.com/health
# → {"status":"ok"}
# Test generation (with user key)
curl -s https://bridge.yourdomain.com/generate-codex \
-H "Authorization: Bearer <USER_KEY>" \
-H "Content-Type: application/json" \
-d '{"systemPrompt":"Reply concisely.","userPrompt":"What is 2+2?"}'
# → {"content":[{"type":"text","text":"4"}],"usage":{...},"cost_usd":0.003}
# Check usage
curl -s https://bridge.yourdomain.com/admin/keys/yourname \
-H "Authorization: Bearer <ADMIN_KEY>"
# → {"name":"yourname","limits":{...},"usage":{"today":{"requests":1,"tokens":9132,"costUsd":0.003},...}}const response = await fetch('https://bridge.yourdomain.com/generate', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': 'Bearer <USER_KEY>',
},
body: JSON.stringify({
systemPrompt: 'You are a helpful assistant.',
userPrompt: 'Explain monads in one sentence.',
model: 'claude-sonnet-4-20250514', // optional
}),
});
const data = await response.json();
console.log(data.content[0].text);The bridge was originally built for a Figma plugin. The provider pattern uses the API key field to pass both the bridge key and URL:
API Key field value: <USER_KEY>@https://bridge.yourdomain.com
The provider code parses this:
function parseBridgeConfig(apiKey: string) {
if (apiKey && apiKey.includes('@')) {
const atIdx = apiKey.indexOf('@');
return {
url: apiKey.slice(atIdx + 1),
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${apiKey.slice(0, atIdx)}`,
},
};
}
return { url: 'http://localhost:3456', headers: { 'Content-Type': 'application/json' } };
}All endpoints return the same shape (Anthropic Messages API compatible):
{
"content": [{ "type": "text", "text": "..." }],
"usage": {
"input_tokens": 1234,
"output_tokens": 56,
"cache_creation_input_tokens": 0,
"cache_read_input_tokens": 789
},
"cost_usd": 0.003,
"duration_ms": 4500
}All user endpoints require Authorization: Bearer <USER_KEY> (unless auth is disabled).
Returns {"status":"ok"}. Unauthenticated — use for uptime monitoring.
Calls Claude Code CLI.
Request:
{
"systemPrompt": "string (required, max 500K chars)",
"userPrompt": "string (required, max 500K chars)",
"model": "string (optional, default: claude-sonnet-4-20250514)"
}Response: See response shape above.
Calls Codex CLI.
Request: Same as /generate. Default model: gpt-5.3-codex.
Response: Same shape. cost_usd is estimated from a hardcoded pricing table.
All admin endpoints require Authorization: Bearer <ADMIN_KEY>.
List all keys with limits and current usage.
{
"keys": [
{
"name": "alice",
"createdAt": "2026-02-15T13:42:45.911Z",
"limits": {
"maxRequestsPerDay": 0,
"maxRequestsPerMonth": 0,
"maxTokensPerMonth": 0,
"maxCostPerDay": 0,
"maxCostPerMonth": 5.00
},
"usage": {
"today": { "requests": 3, "tokens": 27382, "costUsd": 0.009 },
"thisMonth": { "requests": 45, "tokens": 412000, "costUsd": 1.23 }
}
}
]
}Create a new user key.
Request:
{
"name": "string (required, unique)",
"maxRequestsPerDay": 0,
"maxRequestsPerMonth": 0,
"maxTokensPerMonth": 0,
"maxCostPerDay": 0,
"maxCostPerMonth": 0
}All limit fields are optional (default 0 = unlimited).
Response (201):
{
"message": "Key created for \"alice\"",
"key": "abc123...",
"note": "Save this key — it cannot be retrieved again."
}Get a specific key's limits and usage.
Update a key's limits. Only include fields you want to change.
{ "maxCostPerMonth": 10.00 }Revoke a key. Deletes the key and all associated usage data.
Reset a key's usage counters to zero.
| Status | Meaning |
|---|---|
| 400 | Invalid request body (missing/wrong types) |
| 401 | Missing or invalid Authorization header |
| 403 | Invalid key |
| 409 | Key name already exists (on create) |
| 429 | Rate limit or per-key usage limit exceeded |
| 500 | Generation failed (CLI error) |
Limit error messages include the limit that was hit:
{ "error": "Daily cost limit reached ($1.00/day)" }
{ "error": "Monthly request limit reached (100/month)" }From your local machine:
# After making changes locally
bun run build # Verify it compiles
# Deploy
rsync -avz --exclude node_modules --exclude .env --exclude data \
./ai-cli-bridge/ ai-bridge:/opt/ai-cli-bridge/
# Restart on the droplet
ssh ai-bridge 'cd /opt/ai-cli-bridge && bun install --frozen-lockfile && systemctl restart ai-cli-bridge'Subscription tokens expire periodically. When generation requests start failing:
ssh ai-bridge
claude # Re-authenticate Claude
codex auth # Re-authenticate Codex
# No server restart needed — the CLIs read fresh tokens on each invocation# View real-time logs
ssh ai-bridge 'journalctl -u ai-cli-bridge -f'
# Check server status
ssh ai-bridge 'systemctl status ai-cli-bridge'
# Check tunnel status
ssh ai-bridge 'systemctl status cloudflared'
# Check all keys' usage
curl -s https://bridge.yourdomain.com/admin/keys \
-H "Authorization: Bearer <ADMIN_KEY>"# Delete the compromised key
curl -X DELETE https://bridge.yourdomain.com/admin/keys/compromised-user \
-H "Authorization: Bearer <ADMIN_KEY>"
# Create a replacement
curl -X POST https://bridge.yourdomain.com/admin/keys \
-H "Authorization: Bearer <ADMIN_KEY>" \
-H "Content-Type: application/json" \
-d '{"name":"compromised-user","maxCostPerMonth":5.00}'
# Give the new key to the userWhile per-request cost is "free" (covered by your subscription), be aware:
- Claude Max has monthly usage limits that vary by tier
- OpenAI Pro has similar caps
- Per-key cost tracking lets you monitor spending per user
- Use
maxCostPerDay/maxCostPerMonthto cap individual users - Heavy automated usage may hit subscription throttling before the month ends
- Create
src/providers/newprovider.tsfollowing the existing pattern - Export a
generateWithNewProvider(req, cfg): Promise<Response>function - Add the route in
src/server.ts - Add config entries in
src/config.tsand.env.example - Install the CLI on the droplet
Replace execFile() with spawn() and pipe chunks to an SSE response:
import { spawn } from 'child_process';
const child = spawn('claude', args);
res.setHeader('Content-Type', 'text/event-stream');
child.stdout.on('data', chunk => res.write(`data: ${chunk}\n\n`));
child.on('close', () => res.end());