|
| 1 | +# Etappe 44 — The calls page reports on a counter that resets |
| 2 | + |
| 3 | +P2-29, reported by the operator 2026-08-16: |
| 4 | + |
| 5 | +> "calls soll auch wenn möglich audit zeigen verbindungen zeigen maby auch |
| 6 | +> statistik (statistik maby ausweiten auf das ganze) / logs länge" |
| 7 | +
|
| 8 | +Four asks: per-call audit, live connections, RTC statistics, and a wider statistics |
| 9 | +story. The backlog entry said this needed an inventory pass before any of it, because |
| 10 | +§4.24 is the standing warning here — the calls page once reported confidently on the |
| 11 | +SFU while the failing calls were legacy 1:1, which it had no way to see. |
| 12 | + |
| 13 | +## Inventory |
| 14 | + |
| 15 | +**What LiveKit exposes on this deployment**, read from the SFU's own metrics port: |
| 16 | + |
| 17 | +| metric | kind | what it answers | |
| 18 | +|---|---|---| |
| 19 | +| `livekit_room_total` | gauge | rooms open **right now** | |
| 20 | +| `livekit_participant_total` | gauge | participants **right now** | |
| 21 | +| `livekit_room_duration_seconds_count` | counter | rooms that have completed | |
| 22 | +| `livekit_room_duration_seconds_sum` | counter | seconds of room time | |
| 23 | +| `livekit_quality_score_{count,sum}` | histogram | call-quality samples | |
| 24 | +| `livekit_forward_latency_ns_count` | counter | forwarded-media samples | |
| 25 | +| `livekit_node_packet_total` | counter | packets, rises without any call | |
| 26 | + |
| 27 | +Three of these are already parsed — `internal/rtc/media.go` reads the counters that |
| 28 | +prove media flowed (E23). The two **gauges** are not read by anything, and they are |
| 29 | +exactly "Verbindungen zeigen". |
| 30 | + |
| 31 | +**The fact that shapes everything else:** all of these are *process-lifetime*. Read |
| 32 | +live, ten hours after the 26.8.0 upgrade, every single one is `0`: |
| 33 | + |
| 34 | + livekit_room_total 0 |
| 35 | + livekit_participant_total 0 |
| 36 | + livekit_room_duration_seconds_count 0 |
| 37 | + livekit_quality_score_count 0 |
| 38 | + |
| 39 | +That is not "this server has never carried a call". It is "this **process** has not", |
| 40 | +and the process is younger than the counters suggest, because the post-upgrade hook |
| 41 | +deletes the SFU pod on every ESS upgrade to restore `hostNetwork`. On this instance |
| 42 | +that is several times a week. |
| 43 | + |
| 44 | +So a statistics page built directly on these numbers would silently mean "since the |
| 45 | +last upgrade" while looking like "ever" — the §4.24 mistake with a different subject. |
| 46 | +**History has to be recorded by MatrixCtrl or it does not exist.** |
| 47 | + |
| 48 | +**The precedent is already in the package.** `internal/rtc/watcher.go` + `store.go` |
| 49 | +exist for exactly this reason, for a different quantity: DNS observations are sampled |
| 50 | +on a timer and persisted, because "a history built from page views has gaps exactly |
| 51 | +where nobody was looking, which is most of the time." The same sentence applies here. |
| 52 | + |
| 53 | +**What is deliberately not built:** LiveKit's RoomService API (port 7880) would give |
| 54 | +room names and participant identities. It needs the SFU's API secret and a minted |
| 55 | +admin token, and what it returns is *who is in a call with whom*. That is a different |
| 56 | +class of data from "three people are in a call", and it is not needed to answer any |
| 57 | +of the four asks. The gauges answer "connections" without it. |
| 58 | + |
| 59 | +## What this etappe does |
| 60 | + |
| 61 | +**Sampling.** A `Sampler` alongside the existing `Watcher`: reads the metrics port on |
| 62 | +a timer, writes one row per observation. The endpoint is already fetched by the RTC |
| 63 | +handler, so this is a second caller of a known-good path, not a new integration. |
| 64 | + |
| 65 | +**Counter resets are first-class.** When a counter comes back *lower* than the last |
| 66 | +sample, the SFU restarted. The delta is then the new value, not a negative number — |
| 67 | +and the restart itself is worth recording, because "the SFU restarted" explains a gap |
| 68 | +in every other series on the page. |
| 69 | + |
| 70 | +**Exact totals, coarse timing.** `room_duration_seconds_{count,sum}` are cumulative, |
| 71 | +so the number of calls and the total minutes between two samples are exact *even for |
| 72 | +calls that began and ended entirely between them*. What the sampling interval bounds |
| 73 | +is only *when* — the same trade the Watcher documents for DNS, and it gets stated on |
| 74 | +the page rather than left for someone to discover. |
| 75 | + |
| 76 | +**The page gains two sections:** what is happening now (rooms, participants, and how |
| 77 | +long the SFU has been up so the zero can be read correctly), and what has happened |
| 78 | +(calls per day, total minutes, quality samples, restarts). The existing call-path |
| 79 | +assessment stays exactly where it is — it answers a question neither of these does. |
| 80 | + |
| 81 | +**"Statistik ausweiten auf das ganze" is not in scope**, and the plan says so rather |
| 82 | +than half-doing it. A panel-wide statistics story is a design question about what an |
| 83 | +operator wants to see over time, not a matter of finding more counters. This etappe |
| 84 | +makes one series real; the shape of the general case should be argued from a working |
| 85 | +example rather than in advance. |
| 86 | + |
| 87 | +## Definition of done |
| 88 | + |
| 89 | +- A sample is written on a timer and survives an SFU restart |
| 90 | +- A restart is visible as a restart, never as a negative delta or a reset to zero |
| 91 | +- The live section distinguishes "no calls" from "the SFU restarted a minute ago" |
| 92 | +- Totals are exact across a restart, verified against a counter that actually moved |
| 93 | +- The page states the interval, so "when" is never read as more precise than it is |
| 94 | +- No participant identity is read or stored |
| 95 | +- `make check` green, and a live test against the real SFU |
0 commit comments