A Claude Code skill that routes every prompt and tool output through a local vLLM model on a GKE Autopilot Spot GPU — reducing tokens sent to Claude, scrubbing PII, compressing large outputs, and answering simple questions locally before Claude ever sees them.
| Hook | Trigger | Action |
|---|---|---|
UserPromptSubmit |
Every prompt | Classify, optimize for token efficiency, scrub PII, answer factual questions locally |
PreToolUse (Bash) |
Verbose commands | Inject a hint telling Claude what to focus on in the output |
PostToolUse (Bash) |
Large bash output | Compress npm/pytest/kubectl/docker output, keep errors and key results |
PostToolUse (WebFetch) |
Web pages | Compress to code snippets, API details, concrete facts |
SessionStart |
Claude opens | Scale up GKE pod, try L4 Spot → T4 Spot automatically |
Stop |
Claude closes | Kill port-forward |
| Cron (15 min) | No Claude running | Scale pod to 0 — stop paying |
Token savings show live in the Claude Code status bar.
- GKE Autopilot cluster with Spot GPU node pool
kubectlconfigured and pointing at your clusterjq,curl,shasum(standard on macOS/Linux)- A Hugging Face account (for model download — no token needed for public models)
git clone https://github.com/YOUR_USERNAME/gpu-optimizer
cd gpu-optimizer
./install.shThe installer asks for your cluster config (namespace, deployment name, etc.) and handles everything else: writes config, installs hooks, patches ~/.claude/settings.json, sets up cron.
Then apply the Kubernetes manifest:
kubectl apply -f ~/.claude/gpu-optimizer/llm-inference.yamlRestart Claude Code — the status bar shows ✓ L4 vLLM or ✓ T4 vLLM when the pod is ready.
The manifest deploys Qwen2.5-14B-Instruct-AWQ on an nvidia-l4 Spot node. The GPU scheduler automatically falls back to nvidia-tesla-t4 Spot if L4 capacity isn't available in your zone. Both work — the AWQ-quantized model fits on the T4's 16GB VRAM (weights load at ~9.4 GiB).
The hooks automatically detect which GPU is active and tune max_tokens accordingly to stay within the 20s hook timeout.
Approximate cost: $0.19–0.30/hr while the pod is running (T4 Spot GPU + node overhead). The pod scales to 0 automatically when you close Claude, so you only pay while working.
Different model: edit ~/.claude/gpu-optimizer/llm-inference.yaml, update VLLM_MODEL in ~/.claude/gpu-optimizer.conf, apply the manifest, delete the running pod.
Different GPU: change cloud.google.com/gke-accelerator in the manifest and update the node memory request accordingly.
Tune behavior: edit ~/.claude/gpu-optimizer.conf — all hook parameters are configurable without touching the scripts.
Once installed, /gpu-optimizer is available in Claude Code:
/gpu-optimizer status — pod status, vLLM health, active GPU
/gpu-optimizer logs — recent pod logs
/gpu-optimizer savings — token savings breakdown by prompt type
/gpu-optimizer scale up — manually start the pod
/gpu-optimizer cache clear — flush the prompt cache
Every prompt goes through a local classification pipeline before Claude sees it:
- PII scrub — emails, SSNs, credit card numbers replaced with
[EMAIL]etc. - Cache check — SHA256 keyed, 1hr TTL. Repeated prompts cost nothing.
- vLLM classify — type (coding/debugging/refactor/question/creative), language, optimal compression mode
- Optimize — rewrite for token efficiency (conciseness mode for simple prompts, performance mode for complex ones)
- Local answer — for factual questions with no codebase context, the GPU answers directly and Claude echoes it verbatim
The status bar shows: 🔧 ✓ L4 vLLM | 12 optimized | 847 saved | last [p/coding]: 94→31
~/.claude/
├── gpu-optimizer.conf # your config (auto-generated by install.sh)
├── gpu-optimizer/
│ └── llm-inference.yaml # k8s manifest (copy of k8s/ from this repo)
├── gpu-active.json # current GPU type and zone
├── gpu-optimizer-stats.json # session savings tracking
├── gpu-optimizer-cache/ # prompt cache
└── hooks/
├── prompt-optimizer.sh # UserPromptSubmit
├── bash-pre.sh # PreToolUse/Bash
├── bash-compressor.sh # PostToolUse/Bash
├── fetch-compressor.sh # PostToolUse/WebFetch
├── gpu-scheduler.sh # SessionStart
├── optimizer-status.sh # StatusLine
└── llm-autoscale-down.sh # Cron
# Remove hooks
rm ~/.claude/hooks/{prompt-optimizer,bash-pre,bash-compressor,fetch-compressor,gpu-scheduler,optimizer-status,llm-autoscale-down}.sh
# Remove config and assets
rm -rf ~/.claude/gpu-optimizer* ~/.claude/gpu-active.json
# Remove cron
crontab -l | grep -v llm-autoscale-down | crontab -
# Restore settings.json backup
cp ~/.claude/settings.json.bak ~/.claude/settings.json
# Tear down the k8s resources
kubectl delete deployment llm-inference -n default
kubectl delete svc llm-inference -n default
kubectl delete pvc model-cache -n default