Goal
Bring first-audio and barge-in p95 closer to the release targets under a long-running local GPU profile.
Scope
- Separate cached deterministic responses from normal LLM/TTS responses.
- Measure first audio p50/p95 and underruns over a longer soak.
- Review speculative turn-taking thresholds and TTS first-clause budget.
- Track GPU headroom and thermal behavior.
Acceptance
latency_report.py output clearly separates response classes.
- p95 regressions are visible in CI/manual release evidence.
- Any changed threshold is justified by measured behavior.
Goal
Bring first-audio and barge-in p95 closer to the release targets under a long-running local GPU profile.
Scope
Acceptance
latency_report.pyoutput clearly separates response classes.