What krm-foyer limits on its own, why, and what it deliberately leaves alone. The design lists the bounds as requirements; this document explains each choice, records what was left out and why, and says how to read the metrics that show how close real traffic comes to each bound.
A bound is something krm-foyer limits, such as how long a response may stay open. Its limit is the configured value, such as 30 minutes.
Decisions (2026-10-02):
- Native watches stay open, and are not a special case. Every bound applies to every
request through
/k8s, so krm-foyer never has to tell a watch from any other request. Live-view recovery belongs in a client library (watches); the gateway connector ships today and a native connector is requested. See native watches. - Five bounds, each configurable with a documented default: the request rate per session, concurrent requests per session and per replica, how long a response may stay open, and the bytes of a response, counted decoded. On top of them, every open response ends when its session ends.
- A response cut short is aborted, never ended cleanly, so a part of an answer can never pass for the whole of it. See cutting a response short.
- A response too large gets a 502 only when its size is known before it starts. Otherwise the answer has begun, and it is cut short. See response bytes.
- No bound on page size. It would not reliably bound what a list costs. See left out.
- Streams are bounded on their own (2026-10-02): how many a session, and a replica,
may have open, counted apart from requests through
/k8s. Opening one draws on the same request rate, and the response duration and the session check hold it like any other response. See streams. - Metrics come from the Prometheus client, on a listener of their own, labelled only by fixed sets of values. OpenTelemetry is kept for tracing, if that is ever added. See metrics.
krm-foyer is for applications that talk to Kubernetes from a browser, so the bounds are shaped by what a browser does: what a page sends when it loads, what a tab keeps open, and what goes wrong in page code. They protect krm-foyer and the API server from load that no person generates: a retry loop, a leaked watch, a component that turns every render into a request, a compressed answer that expands a thousandfold.
They are not three other things:
- Not an access rule. A bound never changes what a request may do, only how much and how long. Kubernetes alone decides what is allowed.
- Not a limit on export. A user allowed to list a namespace can retrieve all of it through repeated pages, and should be able to.
- Not a defense against a determined user. The per-session bounds are per session, and a new session costs one login. A signed-in user who wants to cause more load can open more sessions, or sign in at the issuer and call the API server directly. Load from a user is the API server's to limit: API Priority and Fairness charges every request to the user who made it, whichever client sent it.
Every bound is per replica. krm-foyer keeps no shared counters. With one replica, as today, that is the whole deployment. Once there are several (shared storage on the roadmap), a session whose requests reach N replicas gets up to N times its per-session limits, and the per-replica limits are what hold.
The per-replica bounds on concurrent requests are the ones that protect krm-foyer itself, however many sessions there are. The byte bound limits one response, not the total: krm-foyer streams, so a response costs it the same memory whatever its size, and the number of responses in flight is what the concurrency bound caps.
A native watch is a GET that the API server answers with a stream of events. A browser
can read one: fetch exposes a response body as a stream, so page code can read events
as they arrive. The code must then do what every Kubernetes client does: parse events,
complete initialization before resuming from a checkpoint, reconnect when the stream
ends, and start a fresh snapshot when history expires.
krm-stream's gateway and browser library handle live-view recovery for applications.
They are the current browser recommendation (watches), and the hello
example uses them. So the question was whether native watches through /k8s should stay at
all. They stay, for three reasons:
- Refusing them would be an access rule. RBAC already says who may watch what. krm-foyer refusing a verb that RBAC allows is exactly the second set of rules the design avoids. krm-foyer blocks as little as possible.
- They work today, against the real API server, and a spec proves it.
- They cost almost nothing extra. What they need, an end when the session ends and cancellation at the API server, every long response needs too: a followed log, or a slow list open at logout. krm-stream's per-user streams use the same gate.
There is no watch client in the browser helper, or an example that reads native watches. The requested native connector belongs in krm-stream and should reuse its browser primitives. It would make native access convenient without duplicating recovery in foyer or individual pages.
krm-foyer does not tell watches apart. It was going to, and two measurements against the real API server (k3s in the e2e fixture, 2026-10-02) showed why it should not:
- The response does not say. A JSON watch is served as plain
application/json; only the other encodings get;stream=watch. That is in Kubernetes' watch handler. - The request has many spellings. The API server takes the first
watchvalue after Go's query parsing: anything but0orfalse, in any case, is a watch, including an empty value,watch=yesand a barewatch.%77atch=1is a watch, a pair with a semicolon is dropped, and the deprecated/watch/paths are watches too. AGETof one named object ignoreswatch. Matching all of that is a second parser of exactly the kind krm-foyer avoids for paths.
So the bounds are written not to need it. Concurrency counts every request in flight, the duration and the session check apply to every response, and a response cut short is aborted whatever it is. A watch is simply a request that stays open.
krm-stream's streams are bounded on their own: see streams.
| Bound | Flag | Default | When it is reached |
|---|---|---|---|
| Request rate per session | -session-request-rate, -session-request-burst |
20 a second, bursts of 100 | 429 RequestRateExceeded with Retry-After |
| Concurrent requests per session | -max-session-concurrent-requests |
64 | 429 TooManyConcurrentRequests |
| Concurrent requests per replica | -max-concurrent-requests |
2000 | 429 TooManyConcurrentRequests |
| Streams open per session | -max-session-streams |
32 | 429 TooManyStreams |
| Streams open per replica | -max-streams |
2000 | 429 TooManyStreams |
| Response duration | -max-response-duration |
30 minutes | The response is cut short |
| Response bytes, decoded | -max-response-bytes |
32 MiB | 502 ResponseTooLarge if known in advance; otherwise cut short |
| Session check | -session-check-interval |
5 seconds | An open response whose session has ended is cut short |
| One write of a stream | -stream-write-timeout |
10 seconds | The stream ends: its browser stopped reading |
Why these defaults:
- 20 requests a second, bursts of 100. The burst covers a cold page load, discovery included, which can send dozens of requests at once. Twenty a second, sustained, is far above what a person clicking generates, and a tight loop meets the bound within seconds.
- 64 concurrent requests per session. Over HTTP/1.1 a browser opens at most six connections to one host, so it never gets near this. Over HTTP/2, a cold page load can have dozens of requests in flight for a moment. Sixty-four leaves room for that next to a dozen open watches. A page that leaks watches meets the limit, and from then on every request of that session is refused, which makes the leak hard to miss.
- 2000 concurrent requests per replica. Each one holds a connection or HTTP/2 stream
to the browser and another to the API server. This is still a capacity hypothesis:
the rehearsal measured streams, not native watches through
/k8s, which pass through a lighter path. Measure them before relying on it. - 32 streams per session. A page shows a handful of live lists, and one session is every tab of the application in that browser. Thirty-two leaves room for several tabs with several streams each, and a page that leaks streams meets it.
- 2000 streams per replica, provisionally. In the rehearsal, 1800 streams of one small note each took about 250 MiB of resident memory on one replica, so 2000 like them would take about 300 MiB. That holds for that workload only: a stream's cost grows with its snapshot, its objects and how often they change. Measure your own scopes before sizing a replica by it.
- 30 minutes per response. Ordinary requests finish in seconds; the API server gives
them 60 at most. What stays open is a watch or a followed log. Thirty minutes is the
shortest time the API server itself keeps a watch open when the client names no
timeoutSeconds(it picks between--min-request-timeout, 1800 seconds by default, and twice that). Clients that name a shortertimeoutSeconds, as client-go does (five to ten minutes), never meet this bound. A session's end is fixed at login, so neither an open response nor a reconnect keeps it, or ends it sooner. - 32 MiB per response. That fits an unpaginated list of a few thousand ordinary objects. gzip expands at most about a thousandfold, so a compressed bomb is stopped after some 32 KiB of what the API server sent.
- A session check every 5 seconds. Logout and expiry cut open responses short within that time, or twice it at worst (see the session check).
A refusal is an interruption, like every other answer
krm-foyer gives instead of the API server's. It carries the Krm-Foyer-Interruption
header, and its Status names the bound in details.causes: one cause with reason
BoundReached, the bound as the metrics name it in field (such as
session_concurrent_requests), and its limit in message.
That header is what tells krm-foyer's 429 from the API server's. The API server sends 429s of its own, from Priority and Fairness, and they pass through unchanged and without the header, since the proxy passes no upstream header outside its allowlist.
The header says that krm-foyer answered. It does not, on its own, say that Kubernetes never saw the request; that depends on the reason, and the interruptions table says it for each:
RequestRateExceeded,TooManyConcurrentRequestsandTooManyStreamsare decided before anything is sent. The request did not reach Kubernetes, so code may send it again, a change included.ResponseTooLargeis decided on the API server's answer. The request reached Kubernetes and may have taken effect. Code may repeat aGET, but must never resend a change on this answer without finding out first whether it happened.
Retry-After comes only where waiting is known to help: for the request rate it is the
time until the session may send its next request, rounded up to whole seconds. A
concurrency slot frees when another request ends, and nobody knows when that will be, so
TooManyConcurrentRequests carries no Retry-After rather than a number made up.
A response can end in five ways:
| Cause | What the browser sees |
|---|---|
The API server ends it: the answer is complete, or a watch reached its timeoutSeconds |
The end of the response |
| The browser goes away | Nothing; nobody is left |
| The response duration is reached | An aborted response |
| The response-byte bound is reached | An aborted response |
| Its session ends: logout or a new login in this process, or absolute or token expiry | An aborted response, within the session check |
When krm-foyer cuts a response short, it does two things:
- It cancels its request to the API server, so the API server releases the watch, or stops the work, at once rather than at its own timeout.
- It aborts the response to the browser. Over HTTP/2 the stream is reset; over HTTP/1.1 the connection is closed before the final chunk. The browser sees a network error, never a clean end.
A clean end would be friendlier to a watch, but krm-foyer does not know which responses
are watches, and for anything else a clean end would make the first part of an answer
look like the whole of it. For a watch the abort makes no difference: Kubernetes clients
resume after a broken connection exactly as after an end, from the last
resourceVersion they saw. If the session has ended, the resumed request gets the 401
interruption, and the application signs in again.
Both steps must work while the response is blocked. A read from the API server is unblocked by cancelling its request. A write to a browser that has stopped reading is unblocked by a write deadline, so a stalled tab cannot keep a response open past its bound.
Every open response asks, once per interval, whether its session is still live: not past its end, and not ended in this process since the response opened, by a logout or a new login in the same browser. The answer needs no storage: the end came with the cookie, and a logout marks the responses a small registry holds while they are open. A copy of the cookie that opens a response after the logout is not marked: logout revokes nothing (sessions).
The check is given at most one interval to answer, and one that does not counts as a session that has ended: fail closed, as the session lifecycle requires. So a response is cut short within one interval of its session ending, plus at most one more for a check that hangs: 5 seconds normally, 10 at worst, with the defaults.
The check reads nothing: it compares two times and a flag, in memory. More replicas would need logout to reach the others' registries, a design the roadmap leaves for when they are needed.
The bound counts decoded bytes. krm-foyer asks the API server for gzip and decodes it, so the bound applies to what the browser receives, and a small compressed body cannot expand past it.
The design first said a response over the bound gets a 502. That can only be true before the response starts, and usually it has started already. krm-foyer streams: it passes on the status and headers as soon as they arrive, then the body as it comes. The API server compresses large answers, so their decoded size is unknown until every byte has been counted, and by then the 200 is on its way to the browser.
The exact rule, for a limit of L bytes:
- A response of up to L decoded bytes passes, whether its length was known in advance or not.
- Known in advance: a response whose
Content-Lengthis more than L gets a 502ResponseTooLargebefore any of it is sent. - Counted as it streams: when byte L + 1 arrives, it is not forwarded, and the response is cut short.
- A
HEADresponse is never too large: it has no body, and itsContent-Lengthdescribes theGET.
The tests send L − 1, L and L + 1 bytes, both with a known length and gzip-encoded.
Other ways to keep the 502, and why not:
- Buffer the response until it is complete or over the limit. Then the status could still change. But memory becomes the limit times the requests in flight, 32 MiB a hundred times over is more than 3 GiB, and a bound meant to protect krm-foyer would become the easiest way to exhaust it.
- Ask the API server not to compress. Large lists would then cross the network several times larger, and a chunked response still has no length in advance.
- Report the overflow in a trailer.
fetchdoes not expose trailers, so no browser code would see it.
Revisit when the held-back page of the interruptions table learns to show a response as escaped text. It needs a small buffer for that page anyway, and a 502 for small bodies could share it.
A stream on /stream/v1 is krm-stream's: a snapshot of a resource, then its changes, for
as long as the page keeps it open (design). krm-foyer
counts streams apart from requests, and counts the watches they hold at the API server
apart again:
-
Streams have limits of their own, per session and per replica. A page that keeps a few streams open should not lose its room for ordinary requests, and a request should not be refused because live views are open. Past a limit, a stream is a 429
TooManyStreamsinterruption, and krm-stream's browser client tries again later. -
Opening a stream is a request for the request rate: one budget per session for
/k8sand/stream, so a page that reconnects in a loop meets the rate bound. A stream that fails in a way that may pass closes, and krm-stream's browser client opens it again after a wait, so its retries are requests too, and meet the same bound. -
The response duration and the session check hold a stream as they hold every response: it is cut short at 30 minutes, and krm-stream's client opens it again with a fresh snapshot, with the same session; and it is cut short within the session-check interval of its session ending.
-
The response-byte bound does not apply. It bounds one answer, and a stream is not one: it is many snapshots and changes over half an hour. The size of a snapshot is the size of a list, which the API server answers the same way through
/k8s. -
Upstream watches are counted, not bounded. A per-user stream holds one watch at the API server, opened as its user; a shared stream holds none of its own, and its scope's shared watch stays open only while some stream reads it. So there are never more upstream watches than streams, and the stream limits bound them. A bound of their own was planned for shared watches, and left out (2026-10-03) for that reason: it could only be reached by first reaching a stream limit. They are counted apart, by whose identity opened them (
krm_foyer_upstream_watches_open{identity}), because with sharing the two counts differ: the reuse is their ratio. -
Each write of a stream is bounded (
-stream-write-timeout, 10 seconds). A tab that stops reading ends its stream within that once the buffers between are full, rather than holding its slot until the response duration. It matters most for a shared stream: krm-stream delivers and rechecks one at a time, so a blocked write would also hold off the recheck that applies a revoked grant. A test stops reading, fills the buffers, revokes the grant and sees the stream end. -
Sessions cost krm-foyer nothing to hold. Each is in its own cookie, so there is no count to bound; what a session can do here is bounded per session above, and its cookie by size, at 3,800 bytes of value (sessions). Sealing one at login and opening one per request cost microseconds (
BenchmarkSealAndOpenininternal/session); the rehearsal records the cookie's size with Dex's tokens.
What is not bounded yet:
- A native watch through
/k8swhose tab stops reading has no separate write timeout; its session ending or response duration still cuts it short. - Snapshot size, per-stream object bookkeeping and shared-cache bytes have no foyer
limit.
-max-response-bytesapplies only to/k8s; it does not protect/stream. A few large scopes can therefore exhaust memory below the stream-count limit. - There is no rate limit on anonymous login/callback traffic. A login in progress is
held by the browser, so starting many costs krm-foyer no memory and locks no one out,
but each callback with a valid login costs a token request to the issuer. The API
gate's per-session rate applies to
/k8sand/stream, not/auth.
These are capacity/admission questions for the next work, not guarantees established by the small-snapshot rehearsal.
The rehearsal (rehearsal_test.go) is part of the e2e
suite and runs with every task verify. It restarts krm-foyer, signs in 200 identities
through Dex's login form, opens nine streams for each, 1800 on one replica, near the
2000 it allows by default, and changes one note they all watch. It restarts krm-foyer
with every stream open, and every identity opens its streams again with the session
cookie it had: no new login. Then every identity signs out. It fails unless every stream had the change within 10 seconds, every stream
was aborted within the session-check interval of its logout, and the streams, the
watches at the API server and the goroutines all went back to where they were. It
prints what it measured with its report.
Three runs on 2026-10-02, against the e2e fixture (k3s 1.36.4 in k3d, krm-stream 0.4.0, one replica, the suite and the cluster on one devcontainer host):
| Measured | Result |
|---|---|
| Opening 1800 streams at once, until every snapshot was complete | 0.7 to 1.1 seconds |
One change reaching all 1800, from kubectl patch starting |
under 110 ms, the slowest; the spread from first to slowest under 50 ms |
| Logout to stream aborted, the slowest of 1800 | 4.1 to 4.5 seconds, within the 5-second check |
| Goroutines per open stream | 6 |
| Resident memory | 30 MiB idle, 34 MiB with 200 sessions, 235 to 267 MiB with 1800 streams: about 110 to 130 KiB per stream, for this workload |
| Watches at the API server | one per stream, 1800, all released at logout |
One run each on 2026-10-03, after shared watches came (krm-stream 0.6.0, otherwise the same fixture), now that the rehearsal runs twice: streams of ConfigMaps, each a watch of its user's own, and streams of shared notes, all on one watch:
| Measured | Per-user (ConfigMaps) | Shared (notes) |
|---|---|---|
| Opening 1800 streams at once, until every snapshot was complete | 0.6 seconds | 2.2 seconds: a SelfSubjectReview for each stream, and the access checks, come first |
| One change reaching all 1800, the slowest | 84 ms | 47 ms |
| Logout to stream aborted, the slowest | 4.4 seconds | 1.6 seconds |
| Goroutines per open stream | 6 | 5 |
| Resident memory with 1800 streams | 326 MiB, about 164 KiB per stream | 220 MiB, about 103 KiB per stream |
| Watches at the API server | 1800 | 1 |
| Access checks while opening | none | 200 asked the API server (two SubjectAccessReviews each), 1600 reused |
| Access checks while held for 31 seconds | none | 200 asked, 12.9 reviews a second; 1600 reused |
One run each on 2026-10-05, after sessions moved into sealed cookies (krm-stream 0.7.0, otherwise the same fixture):
| Measured | Per-user (ConfigMaps) | Shared (notes) |
|---|---|---|
| The session cookie, name and value, with Dex's ID token | 1,589 bytes, of the 4,096 browsers keep | 1,589 bytes |
| Reopening all 1800 streams after a restart, with the cookies from before it | 0.7 seconds, no new login | 2.1 seconds, no new login |
| Opening 1800 streams at once | 0.6 seconds | 2.0 seconds |
| One change reaching all 1800, the slowest | 80 ms | 47 ms |
| Logout to stream aborted, the slowest | 4.6 seconds | 3.3 seconds |
| Resident memory with 200 sessions, before any stream | 40 MiB, 35 idle | 39 MiB, 35 idle |
| Resident memory with 1800 streams | 288 MiB, about 141 KiB per stream | 226 MiB, about 106 KiB per stream |
Opening a sealed cookie costs about 6 µs and 10 KB of allocation per request, and
sealing one at login about 9 µs (go test -bench SealAndOpen ./internal/session/, on
the same host), next to the milliseconds a request to the API server takes.
With sharing, the API server holds one watch instead of 1800, and is asked about each
identity once per recheck rather than about each stream: 12.9 reviews a second for 200
identities at the 30-second default, the rate krm-stream's own estimate gives. The
ConfigMap snapshot holds two objects (the namespace's kube-root-ca.crt beside the
one created), the note snapshot one, so the memory columns are not a like-for-like
comparison of the two paths.
Every stream watches the same scope: one small object. So the memory figures describe
that workload, not a stream in general. krm-stream keeps per-object bookkeeping for each
stream, and its list-then-watch fallback holds a whole snapshot while it lists, so
larger scopes, larger objects and frequent changes cost more, possibly much more.
Resident memory is measured in a fresh process, where it only grows over a run this
short; the garbage collector made smaller measurements too noisy to use. What the
rehearsal does not cover: representative scopes and objects, sustained change, native
watches through /k8s, many replicas, many scopes per shared resource, a slow or
distant API server, and the browser's side. A real cluster's
numbers belong next to these when someone measures them.
A Kubernetes list takes limit, the most items to return in one page; the API server
then answers with a continue token for the next page. A bound on page size would refuse
a limit above some maximum. It was in the design and the roadmap until 2026-10-02.
It would not reliably bound what a list costs:
- Only a
limitthe client sends can be checked. A list without one returns every item, so a client that wants everything at once leaveslimitout, and the bound never applies. - Adding a
limitwhere there is none would break clients silently. A client that sends nolimitdoes not expect acontinuetoken; given a partial list, it shows a partial list, and nobody finds out. Rewriting the request also breaks the rule that krm-foyer passes requests on as they came, or answers itself with an interruption. - What a large list costs krm-foyer and the browser is its bytes, and the
response-byte bound holds those, with or without
limit. - The API server already charges for it. Priority and Fairness estimates the work of
a list from its
limitand the number of objects, and charges it to the user.
Telling a list from a single object, so as to require a limit on lists, would mean
parsing resource paths the way the API server does: a second parser, which krm-foyer
avoids everywhere else.
Revisit when an adopter shows a list that the byte bound lets through and that hurts, or
when an application scope brings resource-path parsing for
reasons of its own. Built then, it should refuse a limit above the maximum with a 400,
not a 429: waiting does not make the request smaller.
Decision (2026-10-02): the Prometheus client, scraped. krm-foyer records its metrics
with the Prometheus Go client and serves them on a listener of its own
(-metrics-listen, :9090 by default, /metrics only). The metrics are never served on
the origin, where any page could read them. Keeping the port to the monitoring system
is the deployment's job: a NetworkPolicy that admits only the monitoring system to
it. The Helm chart ships one, which limits the port to networkPolicy.metricsFrom once
that names the monitoring system; left empty, anyone in the cluster may scrape. The e2e
fixture leaves it empty, and reaches the port as admin through the API server.
The other candidate was the OpenTelemetry metrics API, exported through OpenTelemetry's Prometheus exporter, as gitops-reverser does. It instruments the same way and scrapes the same way. It was not chosen:
- It is the Prometheus client plus a layer. The exporter is built on the Prometheus client, so OpenTelemetry comes on top of it: 21 modules in the build, against 10 for the client alone, for the same scrape.
- Its names are translated. The exporter turns dots into underscores, appends unit
suffixes and
_total, and adds scope labels and atarget_infoseries unless told not to. With the client, the name in the code is the name in the scrape, and in the alerts people build on it. - What it would buy is not needed. Its one advantage is export to more than one
backend, and a scrape already reaches them all: an OpenTelemetry Collector scrapes
/metricswith its Prometheus receiver. Nobody has to run a collector for krm-foyer, and a collector that is down never touches a request.
What is left out on purpose:
- No automatic HTTP instrumentation. Middleware of that kind labels requests by
route, and under
/k8sthe path names namespaces and objects: unbounded label values, and a record of what each user looks at, kept outside the audit log. krm-foyer's metrics use fixed label sets: bound, cause, interruption reason, watch identity class, access-check source and result. They never label a series with a user or scope. Details of a single request belong in its log line. - No tracing yet. Tracing is where OpenTelemetry is the right choice, and it would
come with its tracing SDK only, not its metrics. But the API server accepts a W3C
traceparent, and forwarding the browser's would let a page write into the cluster's traces. That changes the request header allowlist, and needs a decision of its own. - Names, labels and buckets are an interface. People build alerts on them, so a test
reads
/metricsand fails when one changes, and changing one is a breaking change in the changelog. Go runtime and process metrics (go_goroutinesshows a leaked response) come with the client.
Names as scraped:
| Metric | Type | Labels | Meaning |
|---|---|---|---|
krm_foyer_bound_limit |
gauge | bound |
The configured limit of each bound, in its own unit (below) |
krm_foyer_bound_usage_ratio |
histogram | bound |
How much of its bound each request or response let through used, from 0 to 1 |
krm_foyer_bound_reached_total |
counter | bound |
Requests refused, and responses cut short, because a bound was reached |
krm_foyer_requests_in_flight |
gauge | Requests through /k8s open now, watches included |
|
krm_foyer_streams_open |
gauge | Streams through /stream open now |
|
krm_foyer_upstream_watches_open |
gauge | identity |
Watches open at the API server now: user, one per per-user stream, or shared, one per scope in use by shared streams |
krm_foyer_shared_subscriptions_open |
gauge | Streams reading from a shared watch now | |
krm_foyer_shared_overflows_total |
counter | Streams that fell behind their shared watch and got a fresh snapshot from its cache | |
krm_foyer_access_checks_total |
counter | source, result |
Access decisions for shared streams: api_server or cache, and allowed, denied or error |
krm_foyer_access_check_duration_seconds |
histogram | How long a decision that asked the API server took (two SubjectAccessReviews) | |
krm_foyer_subject_reviews_total |
counter | result |
SelfSubjectReviews before a shared stream: resolved, refused or error |
krm_foyer_responses_cut_short_total |
counter | cause |
Responses krm-foyer cut short, by why |
krm_foyer_interruptions_total |
counter | reason |
Answers krm-foyer gave instead of the API server's, by the reason in the Krm-Foyer-Interruption header |
The bound label, and when usage is measured:
bound |
Limit, in | Usage is measured | as |
|---|---|---|---|
session_request_rate |
requests a second | (limit only: the rate the burst refills at) | |
session_request_burst |
requests | at every request let through | the share of the session's burst spent, this request included |
session_concurrent_requests |
requests | at every request let through, as it starts | the session's requests in flight, this one included, over the limit |
concurrent_requests |
requests | at every request let through, as it starts | the replica's requests in flight, this one included, over the limit |
session_streams |
streams | at every stream let through, as it opens | the session's streams open, this one included, over the limit |
streams |
streams | at every stream let through, as it opens | the replica's streams open, this one included, over the limit |
response_duration |
seconds | when a response ends | how long it was open, over the limit |
response_bytes |
bytes | when a response ends | its decoded bytes, over the limit |
A refusal for the request rate is counted under session_request_burst: a request is
refused when the session's burst is spent, and the burst refills at the rate.
The cause label of krm_foyer_responses_cut_short_total is response_duration,
response_bytes or session_ended, the rows of
cutting a response short that are krm-foyer's doing. The
first two are also counted in krm_foyer_bound_reached_total; this counter puts them
next to the one cause that is not a bound.
The usage ratio is the point of these metrics. It shows how close real traffic comes to
each bound, in one unit for all of them, and it stays meaningful when a limit changes.
Multiply by krm_foyer_bound_limit for absolute numbers. The buckets are 0.1, 0.25, 0.5,
0.75, 0.9 and 1.
A value of 1 means the whole bound was used, not that anything was refused: the last
concurrency slot taken, the last of the burst spent, or a response cut short at its
limit. Refusals are not observed in the ratio at all. Only
krm_foyer_bound_reached_total says that a bound refused a request or cut a response
short.
Which bounds were reached today, and how often:
sum by (bound) (increase(krm_foyer_bound_reached_total[1d]))
The 99th percentile, over every request let through, of how much of its session's concurrency limit was in use when it started. It describes requests, not sessions: it does not single out the busiest session, or give any session's peak.
histogram_quantile(0.99,
sum by (le) (rate(krm_foyer_bound_usage_ratio_bucket{bound="session_concurrent_requests"}[1d])))
The share of responses that used more than 90% of the byte bound:
1 - sum(rate(krm_foyer_bound_usage_ratio_bucket{bound="response_bytes", le="0.9"}[1d]))
/ sum(rate(krm_foyer_bound_usage_ratio_count{bound="response_bytes"}[1d]))
Why krm-foyer cuts responses short:
sum by (cause) (rate(krm_foyer_responses_cut_short_total[1h]))
Rules of thumb for tuning:
- Raise a limit when real users reach it. Look at the bound's refusals together with its top bucket: a busy 0.9 to 1 bucket with refusals means the limit is in the way of ordinary use.
- Lower a limit when the 99th percentile stays far below it for weeks: the room above it serves only runaway pages.
- Expect
session_endedat logout and at expiry, one for each response the tab had open. Absolute and token expiry, logout and a new login in the same browser can all end a session's responses. One session can own many responses, so compare against open-response counts rather than expecting one cut per logout. Checks that time out are another cause.