Skip to content

Commit 2f5a47c

Browse files
BxnnyGclaude
andcommitted
Etappe 44: die Calls-Zahlen sterben mit dem Pod, den unser eigener Hook löscht
P2-29 verlangte eine Inventur vor dem Bauen. Sie ergab einen Befund, der die Form aller drei Teile bestimmt hat: jeder LiveKit-Zähler ist prozesslebenslang. Zehn Stunden nach dem 26.8.0-Upgrade stand auf einem Server, der seit Monaten läuft, überall 0 — nicht weil nie telefoniert wurde, sondern weil der Post-Upgrade-Hook den SFU-Pod bei jedem ESS-Upgrade löscht, um hostNetwork wiederherzustellen. Mehrmals pro Woche. Eine Statistikseite direkt auf diesen Zahlen hätte "seit dem letzten Upgrade" gesagt und wie "jemals" geklungen — §4.24 mit anderem Gegenstand. Also aufgezeichnet statt gelesen (§4.42). Das Muster stand schon im Paket: internal/rtc/watcher.go misst DNS auf einem Timer, weil eine aus Seitenaufrufen gebaute Historie genau dort Lücken hat, wo niemand hingesehen hat. - Live: Räume und Teilnehmer, beide die ganze Zeit auf dem Metrics-Port und von nichts gelesen — dieselbe Lücke wie P1-10. - Verlauf: Calls, Gesprächszeit, SFU-Neustarts über 24 h und pro Tag. - Ein Zähler, der niedriger zurückkommt, ist ein Neustart und wird als solcher gespeichert, nie als negatives Delta. Alle drei Zähler werden geprüft, weil der Hauptzähler auf dieser Instanz meist auf beiden Seiten null ist. - Summen sind exakt trotz grobem Intervall: die Zähler sind kumulativ, ein Call zwischen zwei Messungen zählt mit. Nur der Zeitpunkt ist aufs Intervall genau, und die Seite sagt das. - Ein fehlgeschlagener Read schreibt nichts. Eine Null wäre eine erfundene Ruhe und beim nächsten Read ein erfundener Neustart. Teilnehmer-Identitäten werden bewusst nicht gelesen: die RoomService-API würde "wer telefoniert mit wem" liefern, und keine Frage dieser Seite braucht das. Live geprüft: alle sechs Metriken auf der echten SFU vorhanden, Sampler schreibt im 60-Sekunden-Takt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
1 parent 9b41093 commit 2f5a47c

17 files changed

Lines changed: 1076 additions & 10 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 33 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -15,6 +15,39 @@ matching image, so a version identifies one exact pair
1515
1616
## [Unreleased]
1717

18+
## [0.1.44] — 2026-08-16
19+
20+
### Added
21+
22+
- **The calls page says whether anyone is actually on a call.** Rooms and
23+
participants, live from the SFU. Both numbers were on the metrics port from the
24+
beginning and nothing read them — which is the same gap P1-10 was about: every
25+
check on this page was green while the feature was completely dead, because none
26+
of them asked whether anyone had ever used it.
27+
28+
- **A call history that survives the SFU.** Calls, talk time and SFU restarts over
29+
24 hours and per day.
30+
31+
This had to be recorded rather than read. Every LiveKit counter is
32+
process-lifetime, and the post-upgrade hook deletes the SFU pod on every ESS
33+
upgrade to restore `hostNetwork` — so a statistics page built directly on those
34+
numbers would silently mean "since the last upgrade" while reading as "ever".
35+
Measured ten hours after the 26.8.0 upgrade, every counter on the server was `0`.
36+
37+
The totals are exact regardless of the sampling interval: the underlying counters
38+
are cumulative, so a call that starts and ends entirely between two samples is
39+
still counted. Only its *timing* is bounded by the interval, and the page says so
40+
rather than leaving it to be discovered.
41+
42+
A counter that comes back lower means the SFU restarted. That is recorded as a
43+
restart — never as a negative delta — and shown, because it explains a
44+
discontinuity in every other series on the page.
45+
46+
- **No participant identities are read or stored.** LiveKit's RoomService API would
47+
give room names and participants; it is deliberately not used. "Three people are
48+
in a call" and "who is in a call with whom" are different classes of data, and
49+
none of the questions this page answers needs the second.
50+
1851
## [0.1.43] — 2026-08-16
1952

2053
Three etappes in one release: 0.1.42 was written but never built — the image build

‎cmd/matrixctrl/main.go‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -266,7 +266,14 @@ func main() {
266266
// Observe the announced RTC address on a timer. Doing it only on page view
267267
// would leave the history with gaps exactly where nobody was looking, and the
268268
// thing being measured is *when* the address changed.
269-
go rtc.NewWatcher(rtc.NewStore(pool), rtcHandler.AnnouncedHost, 0).Start(context.Background())
269+
rtcStore := rtc.NewStore(pool)
270+
go rtc.NewWatcher(rtcStore, rtcHandler.AnnouncedHost, 0).Start(context.Background())
271+
272+
// Sample the SFU's counters on a timer, for a stronger version of the same
273+
// reason: LiveKit's counters are process-lifetime, and the post-upgrade hook
274+
// deletes the SFU pod on every ESS upgrade. Unrecorded, the call history does
275+
// not merely have gaps — it is destroyed several times a week (E44).
276+
go rtc.NewSampler(rtcStore, rtcHandler.MetricsReader(), 0).Start(context.Background())
270277

271278
srv := server.New(addr, router)
272279

‎deploy/helm/matrixctrl/Chart.yaml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2,8 +2,8 @@ apiVersion: v2
22
name: matrixctrl
33
description: Admin layer for self-hosted Matrix / Element Server Suite (ESS)
44
type: application
5-
version: 0.1.43
6-
appVersion: "0.1.43"
5+
version: 0.1.44
6+
appVersion: "0.1.44"
77
keywords:
88
- matrix
99
- element

‎docs/BACKLOG.md‎

Lines changed: 11 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -474,7 +474,17 @@ implemented, OIDC state consumed atomically via `DELETE … RETURNING` (CSRF-saf
474474
*Why it matters:* k3s on ARM boards is a realistic home-server case, and the
475475
README does not currently say the image is amd64-only.
476476

477-
- **P2-29 · Calls show no audit, no connections and no statistics (S14).** Reported
477+
- **P2-29 · Calls show no audit, no connections and no statistics (S14).**
478+
**Two of four done 2026-08-16 (E44, [DESIGN.md §4.42](DESIGN.md)):** live rooms and
479+
participants, and a recorded history of calls, talk time and SFU restarts that
480+
survives the pod. The inventory this entry demanded found the reason none of it
481+
could simply be read: every LiveKit counter is process-lifetime, and the
482+
post-upgrade hook deletes the SFU pod on every ESS upgrade.
483+
*Still open:* per-call audit entries with more than a count — that needs either the
484+
RoomService API (participant identities, deliberately not read) or Synapse-side
485+
call events, and neither is a small addition. And "Statistik ausweiten auf das
486+
ganze", which is a design question rather than a missing counter.
487+
Original entry: Reported
478488
by the operator 2026-08-16: "calls soll auch wenn möglich audit zeigen
479489
verbindungen zeigen maby auch statistik (statistik maby ausweiten auf das ganze) /
480490
logs länge".

‎docs/DESIGN.md‎

Lines changed: 58 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1404,3 +1404,61 @@ waits in one blocking call and the boundary is not observable. `internal/rollout
14041404
keeps its no-client-go rule, so the assembly is a pure function with unit tests, and
14051405
a live test proves the two reads it depends on are permitted under E40's namespaced
14061406
Role and that Generation is actually populated. Affects S11.
1407+
1408+
### §4.42 — A counter that resets is not a history (2026-08-16, operator + agent, etappe 44)
1409+
1410+
The operator asked the calls page for audit, connections and statistics. The
1411+
inventory that P2-29 demanded before building any of it produced one fact that
1412+
decided the shape of all three.
1413+
1414+
Every number LiveKit publishes is **process-lifetime**. Read ten hours after the
1415+
26.8.0 upgrade, on a server that has been running for months:
1416+
1417+
livekit_room_total 0
1418+
livekit_participant_total 0
1419+
livekit_room_duration_seconds_count 0
1420+
livekit_quality_score_count 0
1421+
1422+
None of that means "this deployment has never carried a call". It means "this
1423+
process has not", and the process is young because MatrixCtrl's own post-upgrade
1424+
hook deletes the SFU pod on every ESS upgrade to restore `hostNetwork` — several
1425+
times a week here. A statistics panel reading those counters directly would have
1426+
said "since the last upgrade" in the voice of "ever", which is §4.24 again with a
1427+
different subject: reporting confidently on the part that happens to be readable.
1428+
1429+
**So the history is recorded, not read.** The precedent was already in the package:
1430+
`internal/rtc/watcher.go` samples DNS on a timer because "a history built from page
1431+
views has gaps exactly where nobody was looking, which is most of the time." Here
1432+
the argument is stronger — unrecorded, this history is not merely gappy, it is
1433+
destroyed on a schedule.
1434+
1435+
Three things make the recording trustworthy:
1436+
1437+
- **A reset is a first-class event, not an anomaly to smooth over.** When a counter
1438+
comes back lower, the delta is the *new value* — everything the old process
1439+
counted was already recorded by samples taken while it ran — and the restart is
1440+
stored, because it explains a discontinuity in every other series. All three
1441+
counters are checked, not one: the headline counter is zero on both sides of most
1442+
restarts on this instance, so watching only it would miss them.
1443+
- **Exact totals from an inexact interval.** Both underlying counters are
1444+
cumulative, so calls and minutes between two samples are right even for a call
1445+
that began and ended between them. Only the *timing* is coarse, and the page says
1446+
which is which instead of implying both are precise.
1447+
- **A failed read records nothing.** An unreachable metrics endpoint is not an SFU
1448+
with zero rooms. Writing a zero would fabricate a quiet period, and would then
1449+
read as a counter reset on the next successful sample — inventing a restart that
1450+
never happened.
1451+
1452+
**What was deliberately not built:** LiveKit's RoomService API would give room names
1453+
and participant identities. "Three people are in a call" and "who is in a call with
1454+
whom" are different classes of data, and no question this page answers needs the
1455+
second. The gauges answer "connections" without reading a secret, minting an admin
1456+
token, or storing anyone's identity.
1457+
1458+
**Also not built:** "Statistik ausweiten auf das ganze". A panel-wide statistics
1459+
story is a design question about what an operator wants to see over time, not a
1460+
matter of finding more counters. The plan says so rather than half-doing it — the
1461+
general shape is better argued from one working example than in advance.
1462+
**Consequences:** migration 011 keeps the raw counter values beside the resolved
1463+
deltas, so a bug in the delta logic can be corrected against the original
1464+
observations rather than having destroyed them. Affects S14.

‎docs/ROADMAP.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -51,6 +51,7 @@ Etappes 1–10 are **reconstructed from `git log`** (39 commits, 2026-05-27 →
5151
| 19 | Calling — the ports that must be forwarded, and an explicit "this half is not checkable from here" | ✅ 2026-08-01 · `v0.1.19` · [plan](plans/etappe-19-calling-reachability.md) |
5252
| 32 | Release Notes auf der Upgrade-Seite + Version aus der Liste übernommen — die andere Hälfte der Pin-Warnung | ✅ 2026-08-05 · `v0.1.33` · [plan](plans/etappe-32-release-notes.md) |
5353
| 33 | OIDC-Init wiederholen statt einmalig aufgeben — ein Neustart vor MAS sperrte den Operator 11 h aus dem eigenen Panel aus | ✅ 2026-08-06 · `v0.1.34` · [plan](plans/etappe-33-oidc-retry.md) |
54+
| 44 | Calls: wer gerade telefoniert, und ein Verlauf — die SFU-Zähler sterben mit dem Pod, den der Post-Upgrade-Hook jedes Mal löscht | 🔄 gebaut 2026-08-16 · `v0.1.44` · [plan](plans/etappe-44-call-history.md) |
5455
| 43 | Das Upgrade-Fenster zeigte eine Uhr statt Fortschritt — dazu Versionen mit Datum und Notes, und der Typecheck, der nie etwas geprüft hat | ✅ 2026-08-16 · `v0.1.43` · [plan](plans/etappe-43-upgrade-progress.md) |
5556
| 42 | Räume verbinden ging und funktionierte dann nicht — der Scope, den E36 bewusst weggelassen hatte; dazu Historie über Neustarts hinweg schnell | ✅ 2026-08-16 · `v0.1.43` · [plan](plans/etappe-42-rooms-connect-and-latency.md) |
5657
| 41 | Räume, zweite Hälfte — Detail, Mitglieder, und die eine Aktion, die man zurücknehmen kann | ✅ 2026-08-16 · `v0.1.43` · [plan](plans/etappe-41-room-detail.md) |
Lines changed: 95 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,95 @@
1+
# Etappe 44 — The calls page reports on a counter that resets
2+
3+
P2-29, reported by the operator 2026-08-16:
4+
5+
> "calls soll auch wenn möglich audit zeigen verbindungen zeigen maby auch
6+
> statistik (statistik maby ausweiten auf das ganze) / logs länge"
7+
8+
Four asks: per-call audit, live connections, RTC statistics, and a wider statistics
9+
story. The backlog entry said this needed an inventory pass before any of it, because
10+
§4.24 is the standing warning here — the calls page once reported confidently on the
11+
SFU while the failing calls were legacy 1:1, which it had no way to see.
12+
13+
## Inventory
14+
15+
**What LiveKit exposes on this deployment**, read from the SFU's own metrics port:
16+
17+
| metric | kind | what it answers |
18+
|---|---|---|
19+
| `livekit_room_total` | gauge | rooms open **right now** |
20+
| `livekit_participant_total` | gauge | participants **right now** |
21+
| `livekit_room_duration_seconds_count` | counter | rooms that have completed |
22+
| `livekit_room_duration_seconds_sum` | counter | seconds of room time |
23+
| `livekit_quality_score_{count,sum}` | histogram | call-quality samples |
24+
| `livekit_forward_latency_ns_count` | counter | forwarded-media samples |
25+
| `livekit_node_packet_total` | counter | packets, rises without any call |
26+
27+
Three of these are already parsed — `internal/rtc/media.go` reads the counters that
28+
prove media flowed (E23). The two **gauges** are not read by anything, and they are
29+
exactly "Verbindungen zeigen".
30+
31+
**The fact that shapes everything else:** all of these are *process-lifetime*. Read
32+
live, ten hours after the 26.8.0 upgrade, every single one is `0`:
33+
34+
livekit_room_total 0
35+
livekit_participant_total 0
36+
livekit_room_duration_seconds_count 0
37+
livekit_quality_score_count 0
38+
39+
That is not "this server has never carried a call". It is "this **process** has not",
40+
and the process is younger than the counters suggest, because the post-upgrade hook
41+
deletes the SFU pod on every ESS upgrade to restore `hostNetwork`. On this instance
42+
that is several times a week.
43+
44+
So a statistics page built directly on these numbers would silently mean "since the
45+
last upgrade" while looking like "ever" — the §4.24 mistake with a different subject.
46+
**History has to be recorded by MatrixCtrl or it does not exist.**
47+
48+
**The precedent is already in the package.** `internal/rtc/watcher.go` + `store.go`
49+
exist for exactly this reason, for a different quantity: DNS observations are sampled
50+
on a timer and persisted, because "a history built from page views has gaps exactly
51+
where nobody was looking, which is most of the time." The same sentence applies here.
52+
53+
**What is deliberately not built:** LiveKit's RoomService API (port 7880) would give
54+
room names and participant identities. It needs the SFU's API secret and a minted
55+
admin token, and what it returns is *who is in a call with whom*. That is a different
56+
class of data from "three people are in a call", and it is not needed to answer any
57+
of the four asks. The gauges answer "connections" without it.
58+
59+
## What this etappe does
60+
61+
**Sampling.** A `Sampler` alongside the existing `Watcher`: reads the metrics port on
62+
a timer, writes one row per observation. The endpoint is already fetched by the RTC
63+
handler, so this is a second caller of a known-good path, not a new integration.
64+
65+
**Counter resets are first-class.** When a counter comes back *lower* than the last
66+
sample, the SFU restarted. The delta is then the new value, not a negative number —
67+
and the restart itself is worth recording, because "the SFU restarted" explains a gap
68+
in every other series on the page.
69+
70+
**Exact totals, coarse timing.** `room_duration_seconds_{count,sum}` are cumulative,
71+
so the number of calls and the total minutes between two samples are exact *even for
72+
calls that began and ended entirely between them*. What the sampling interval bounds
73+
is only *when* — the same trade the Watcher documents for DNS, and it gets stated on
74+
the page rather than left for someone to discover.
75+
76+
**The page gains two sections:** what is happening now (rooms, participants, and how
77+
long the SFU has been up so the zero can be read correctly), and what has happened
78+
(calls per day, total minutes, quality samples, restarts). The existing call-path
79+
assessment stays exactly where it is — it answers a question neither of these does.
80+
81+
**"Statistik ausweiten auf das ganze" is not in scope**, and the plan says so rather
82+
than half-doing it. A panel-wide statistics story is a design question about what an
83+
operator wants to see over time, not a matter of finding more counters. This etappe
84+
makes one series real; the shape of the general case should be argued from a working
85+
example rather than in advance.
86+
87+
## Definition of done
88+
89+
- A sample is written on a timer and survives an SFU restart
90+
- A restart is visible as a restart, never as a negative delta or a reset to zero
91+
- The live section distinguishes "no calls" from "the SFU restarted a minute ago"
92+
- Totals are exact across a restart, verified against a counter that actually moved
93+
- The page states the interval, so "when" is never read as more precise than it is
94+
- No participant identity is read or stored
95+
- `make check` green, and a live test against the real SFU

‎internal/api/handlers/rtc.go‎

Lines changed: 51 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -23,13 +23,13 @@ type RTCHandler struct {
2323
configStore *config.Store
2424
essNS string
2525
essRelease string
26-
addresses *rtc.Store
26+
store *rtc.Store
2727
}
2828

2929
func NewRTCHandler(k8sClient *k8s.Client, cfgStore *config.Store, essNS, essRelease string, db *pgxpool.Pool) *RTCHandler {
3030
return &RTCHandler{
3131
k8s: k8sClient, configStore: cfgStore, essNS: essNS, essRelease: essRelease,
32-
addresses: rtc.NewStore(db),
32+
store: rtc.NewStore(db),
3333
}
3434
}
3535

@@ -83,7 +83,7 @@ func (h *RTCHandler) Status(w http.ResponseWriter, r *http.Request) {
8383
if len(ips) > 0 {
8484
resolved = ips[0]
8585
}
86-
if err := h.addresses.Record(ctx, host, resolved); err != nil {
86+
if err := h.store.Record(ctx, host, resolved); err != nil {
8787
log.Printf("rtc: could not record address observation for %q: %v", host, err)
8888
}
8989

@@ -99,7 +99,12 @@ func (h *RTCHandler) Status(w http.ResponseWriter, r *http.Request) {
9999
"ports": ports,
100100
"freshness": freshness,
101101
"media": media,
102-
"call_paths": paths,
102+
// The uptime was already computed for the findings and thrown away. The page
103+
// needs it in its own right: every counter here is process-lifetime, so
104+
// "0 Räume" a minute after a restart and "0 Räume" all day are the same
105+
// number meaning different things (E44).
106+
"sfu_uptime": uptime,
107+
"call_paths": paths,
103108
"findings": append(
104109
rtc.AssessWithFreshness(ports, host, ips, resolveErr, freshness, freshnessDetail),
105110
rtc.AssessMedia(media, mediaOK, uptime),
@@ -138,7 +143,7 @@ func (h *RTCHandler) freshness(ctx context.Context, host string) (rtc.Freshness,
138143
return rtc.FreshnessUnknown, "Für matrixRTC ist kein Hostname konfiguriert."
139144
}
140145

141-
obs, err := h.addresses.Newest(ctx, host)
146+
obs, err := h.store.Newest(ctx, host)
142147
if err != nil {
143148
return rtc.FreshnessUnknown, "Die Adress-Historie konnte nicht gelesen werden."
144149
}
@@ -332,6 +337,47 @@ func (h *RTCHandler) mediaEvidence(ctx context.Context) (rtc.MediaEvidence, bool
332337
return ev, ok, uptime
333338
}
334339

340+
// History answers "what has actually happened on this SFU" — the question the
341+
// metrics port cannot answer about anything before the current process (E44).
342+
func (h *RTCHandler) History(w http.ResponseWriter, r *http.Request) {
343+
ctx, cancel := context.WithTimeout(r.Context(), 10*time.Second)
344+
defer cancel()
345+
346+
day, err := h.store.SamplesSince(ctx, time.Now().Add(-24*time.Hour), 0)
347+
if err != nil {
348+
Error(w, http.StatusInternalServerError, "could not read samples: "+err.Error())
349+
return
350+
}
351+
daily, err := h.store.Daily(ctx, 30)
352+
if err != nil {
353+
Error(w, http.StatusInternalServerError, "could not aggregate samples: "+err.Error())
354+
return
355+
}
356+
357+
JSON(w, http.StatusOK, map[string]any{
358+
"last_24h": rtc.Sum(day),
359+
"daily": daily,
360+
// Reported so the page can say how precise "when" is. The totals do not
361+
// depend on it — both underlying counters are cumulative, so a call that
362+
// starts and ends between two samples is still counted — but the timing
363+
// within an interval is lost, and a reader should be told which is which.
364+
"interval_seconds": int(rtc.SamplerInterval.Seconds()),
365+
})
366+
}
367+
368+
// MetricsReader exposes the SFU read for the sampler.
369+
//
370+
// The uptime is dropped: it comes from the pod's start time rather than the metrics
371+
// body, and a sample is about what the counters said, not about how the process
372+
// that said it was doing. The page still shows uptime, because there the point is
373+
// to let the reader tell "no calls" from "restarted a minute ago".
374+
func (h *RTCHandler) MetricsReader() func(context.Context) (rtc.MediaEvidence, bool) {
375+
return func(ctx context.Context) (rtc.MediaEvidence, bool) {
376+
ev, ok, _ := h.mediaEvidence(ctx)
377+
return ev, ok
378+
}
379+
}
380+
335381
func formatUptime(d time.Duration) string {
336382
switch {
337383
case d < time.Hour:

‎internal/api/router.go‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -90,6 +90,7 @@ func NewRouter(deps Deps) http.Handler {
9090
}
9191
if deps.RTC != nil {
9292
r.Get("/api/v1/rtc/status", deps.RTC.Status)
93+
r.Get("/api/v1/rtc/history", deps.RTC.History)
9394
r.Get("/api/v1/users", deps.Users.List)
9495
}
9596

0 commit comments

Comments
 (0)