What problem are you solving?
Every enabled stage spawns a fresh claude -p, and on my setup each call costs ~9 s wall-clock — roughly 6 s of that is process/auth overhead, not the model. So the default config (correction only) already adds ~9 s to every prompt, and the full pipeline (correction + translation + enhancement, non-English, mid-session) adds ~36 s. Two of those calls are unnecessary in common cases.
Proposed solution
- Gate translation with a fast language-ID classifier (e.g. fasttext
lid.176, under 1 ms on CPU) instead of running the correction agent purely to detect language. This deletes the correction_only_for_language branch in run_pipeline / runPipeline — translation-only English prompts drop from 1 call to 0.
- Add an exact-match cache for the context-free stages (correction, translation), keyed on normalised prompt text and pre-warmed from the existing audit log. Enhancement must be excluded, since its output depends on conversation context.
What I measured
Three runs per stage, end-to-end claude -p wall-clock: correction 9.4 s, translation 8.7 s, enhancement 8.8 s, summarisation 9.0 s. Overhead (wall minus model time) was ~6 s on every call. Per config: correction-only = 1 call ≈ 9.4 s/prompt; all stages, non-English, mid-session = 4 calls ≈ 36 s/prompt. Each added stage is about +9 s, not a marginal cost.
Gain
- The classifier saves a full ~9 s in any translation config; translation-only + English falls to 0 s.
- The cache saves ~9 s on every repeated prompt in the default (correction-only) config.
Before submitting
What problem are you solving?
Every enabled stage spawns a fresh
claude -p, and on my setup each call costs ~9 s wall-clock — roughly 6 s of that is process/auth overhead, not the model. So the default config (correction only) already adds ~9 s to every prompt, and the full pipeline (correction + translation + enhancement, non-English, mid-session) adds ~36 s. Two of those calls are unnecessary in common cases.Proposed solution
lid.176, under 1 ms on CPU) instead of running the correction agent purely to detect language. This deletes thecorrection_only_for_languagebranch inrun_pipeline/runPipeline— translation-only English prompts drop from 1 call to 0.What I measured
Three runs per stage, end-to-end
claude -pwall-clock: correction 9.4 s, translation 8.7 s, enhancement 8.8 s, summarisation 9.0 s. Overhead (wall minus model time) was ~6 s on every call. Per config: correction-only = 1 call ≈ 9.4 s/prompt; all stages, non-English, mid-session = 4 calls ≈ 36 s/prompt. Each added stage is about +9 s, not a marginal cost.Gain
Before submitting