Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

6 Commits

Folders and files

Repository files navigation

gpu-optimizer

A Claude Code skill that routes every prompt and tool output through a local vLLM model on a GKE Autopilot Spot GPU — reducing tokens sent to Claude, scrubbing PII, compressing large outputs, and answering simple questions locally before Claude ever sees them.

What it does

Hook Trigger Action
UserPromptSubmit Every prompt Classify, optimize for token efficiency, scrub PII, answer factual questions locally
PreToolUse (Bash) Verbose commands Inject a hint telling Claude what to focus on in the output
PostToolUse (Bash) Large bash output Compress npm/pytest/kubectl/docker output, keep errors and key results
PostToolUse (WebFetch) Web pages Compress to code snippets, API details, concrete facts
SessionStart Claude opens Scale up GKE pod, try L4 Spot → T4 Spot automatically
Stop Claude closes Kill port-forward
Cron (15 min) No Claude running Scale pod to 0 — stop paying

Token savings show live in the Claude Code status bar.

Requirements

  • GKE Autopilot cluster with Spot GPU node pool
  • kubectl configured and pointing at your cluster
  • jq, curl, shasum (standard on macOS/Linux)
  • A Hugging Face account (for model download — no token needed for public models)

Quick start

git clone https://github.com/YOUR_USERNAME/gpu-optimizer
cd gpu-optimizer
./install.sh

The installer asks for your cluster config (namespace, deployment name, etc.) and handles everything else: writes config, installs hooks, patches ~/.claude/settings.json, sets up cron.

Then apply the Kubernetes manifest:

kubectl apply -f ~/.claude/gpu-optimizer/llm-inference.yaml

Restart Claude Code — the status bar shows ✓ L4 vLLM or ✓ T4 vLLM when the pod is ready.

GPU and model

The manifest deploys Qwen2.5-14B-Instruct-AWQ on an nvidia-l4 Spot node. The GPU scheduler automatically falls back to nvidia-tesla-t4 Spot if L4 capacity isn't available in your zone. Both work — the AWQ-quantized model fits on the T4's 16GB VRAM (weights load at ~9.4 GiB).

The hooks automatically detect which GPU is active and tune max_tokens accordingly to stay within the 20s hook timeout.

Approximate cost: $0.19–0.30/hr while the pod is running (T4 Spot GPU + node overhead). The pod scales to 0 automatically when you close Claude, so you only pay while working.

Customizing

Different model: edit ~/.claude/gpu-optimizer/llm-inference.yaml, update VLLM_MODEL in ~/.claude/gpu-optimizer.conf, apply the manifest, delete the running pod.

Different GPU: change cloud.google.com/gke-accelerator in the manifest and update the node memory request accordingly.

Tune behavior: edit ~/.claude/gpu-optimizer.conf — all hook parameters are configurable without touching the scripts.

Slash command

Once installed, /gpu-optimizer is available in Claude Code:

/gpu-optimizer status    — pod status, vLLM health, active GPU
/gpu-optimizer logs      — recent pod logs
/gpu-optimizer savings   — token savings breakdown by prompt type
/gpu-optimizer scale up  — manually start the pod
/gpu-optimizer cache clear — flush the prompt cache

How the prompt optimizer works

Every prompt goes through a local classification pipeline before Claude sees it:

  1. PII scrub — emails, SSNs, credit card numbers replaced with [EMAIL] etc.
  2. Cache check — SHA256 keyed, 1hr TTL. Repeated prompts cost nothing.
  3. vLLM classify — type (coding/debugging/refactor/question/creative), language, optimal compression mode
  4. Optimize — rewrite for token efficiency (conciseness mode for simple prompts, performance mode for complex ones)
  5. Local answer — for factual questions with no codebase context, the GPU answers directly and Claude echoes it verbatim

The status bar shows: 🔧 ✓ L4 vLLM | 12 optimized | 847 saved | last [p/coding]: 94→31

Architecture

~/.claude/
├── gpu-optimizer.conf          # your config (auto-generated by install.sh)
├── gpu-optimizer/
│   └── llm-inference.yaml      # k8s manifest (copy of k8s/ from this repo)
├── gpu-active.json             # current GPU type and zone
├── gpu-optimizer-stats.json    # session savings tracking
├── gpu-optimizer-cache/        # prompt cache
└── hooks/
    ├── prompt-optimizer.sh     # UserPromptSubmit
    ├── bash-pre.sh             # PreToolUse/Bash
    ├── bash-compressor.sh      # PostToolUse/Bash
    ├── fetch-compressor.sh     # PostToolUse/WebFetch
    ├── gpu-scheduler.sh        # SessionStart
    ├── optimizer-status.sh     # StatusLine
    └── llm-autoscale-down.sh   # Cron

Uninstall

# Remove hooks
rm ~/.claude/hooks/{prompt-optimizer,bash-pre,bash-compressor,fetch-compressor,gpu-scheduler,optimizer-status,llm-autoscale-down}.sh

# Remove config and assets
rm -rf ~/.claude/gpu-optimizer* ~/.claude/gpu-active.json

# Remove cron
crontab -l | grep -v llm-autoscale-down | crontab -

# Restore settings.json backup
cp ~/.claude/settings.json.bak ~/.claude/settings.json

# Tear down the k8s resources
kubectl delete deployment llm-inference -n default
kubectl delete svc llm-inference -n default
kubectl delete pvc model-cache -n default

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages