The router could learn real context limits instead of trusting declared ones (#13922, #14318, #14135, #14931) #14947
Replies: 1 comment
|
Thanks for the write-up, @behrnt-slatgng, and for the disclosure. The diagnosis matches what we've been fixing one number at a time. #14931 is still open, and it's the clearest case: a On your questions: 1. Is provider-reported 2. Would we take a PR? Yes, with a few conditions, since this changes routing:
The estimator correction factor is a good second step. #14931 is the concrete case to validate it against. On "about 32K served regardless of the card": we haven't measured that across providers ourselves, so I'd keep it as a hypothesis the observed data can confirm or refute, not as a design premise. |
Uh oh!
There was an error while loading. Please reload this page.
Disclosure up front: I build Grunz, a hosted chat + coding agent on open-weight models. I'm not selling anything here; I read the recent context-limit issues and think OmniRoute is in an unusually good position to fix the root cause.
Every context-limit bug here is a declared number that was wrong
Going through the last few weeks:
max_tokensis read as the context window and becomes inputTokenLimit (#14159) #14318:max_tokenswas read as the context window.context_length: 128000.chars/4estimate over-counted about 1 MB of Codex tool schemas 5 to 6 times (257,448 estimated vs. 49,822 reported by the provider), then checked that against a 128K fallback thatcombo-minspread across the whole combo.Each fix corrects a declared number: the catalog, the discovery parser, the fallback. That's the right short-term fix, but the declared number can also be wrong in the other direction. On most hosted OpenAI-compatible endpoints I've checked, the context actually served is about 32K no matter what the model card or
/v1/modelssays. When you go over it, some providers reject the request and some just truncate.OmniRoute already has the data to learn the real limit
A router sees two things no catalog does:
prompt_tokensthe provider actually counted. fix(backend): Context guard rejects Codex CLI at ~50k real tokens: chars/4 estimator over-counts ~1MB of tool schemas 5-6x, checked against a generic 128k fallback limit (model catalog miss → combo-min) #14931's 49,822 came from there. Each success is a lower bound on the real ceiling for that provider and model.Keeping
observed_min_okandobserved_max_rejectedper provider/model/account, and preferring them over catalog values once they exist, would make the guard self-correcting. It would also letcombo-mintake the minimum of observed limits instead of the minimum of guesses. And it gives you a way to catch a real window lower than declared, which no catalog fix can.For the estimator, the same data gives you a correction factor for each request shape (the provider's count divided by your estimate), so a tool-heavy client like Codex stops being over-counted without needing a tokenizer per model.
Questions
prompt_tokensstored anywhere per request today, or only surfaced in usage stats?sourcefield, so/v1/modelscan saycontext_length: 262144 (observed)vs.(catalog)vs.(fallback)?All reactions