Skip to content

Commit 4e421a6

Browse files
BxnnyGclaude
andcommitted
Etappe 76: Ein Installer, der die Fallen kennt — und ein Release, das ankommt
v0.1.70 wurde getaggt und nie veröffentlicht. Der Release-Job ist an seinem eigenen Guard gescheitert ("CHANGELOG must have a section for this version"), korrekt, aber hinter dem Tag: GHCR blieb auf 0.1.69, `helm install` löste brav die neueste *veröffentlichte* Version auf, und der Operator lief auf einer frischen VM in genau den Bug, den 0.1.70 behebt. Gemeldet hatte ich es als ausgeliefert — belegt mit lokalem Cluster, Pod-Logs, Revision 73. Alles wahr, alles auf der falschen Seite der Veröffentlichungsgrenze. Sein Terminal, vier Fehlschläge, keiner davon in MatrixCtrl: 1. Kubernetes cluster unreachable: localhost:8080 → KUBECONFIG nicht gesetzt 2. "kept due to the resource policy" → Neuinstallation auf alter DB 3. admin-password leer → 0.1.70 nie veröffentlicht 4. got: %E2%80%A6 → README-Kommando mit echtem "…" Alle vier waren dokumentiert. Dokumentation zum Zeitpunkt des Fehlers hilft nur, wenn man weiß, wonach man sucht — eine Meldung über Port 8080 schickt niemanden in einen Abschnitt über Kubeconfig-Pfade. scripts/install.sh: install · uninstall · password · status. Prüft der Reihe nach genau das, was hier schiefging, und sagt bei jedem Punkt einen Satz statt eines Exit-Codes. Fragt Hostname und TLS-Terminierung aus (cert-manager / Cloudflare Full / Cloudflare Flexible / kein TLS — vier Kombinationen aus entrypoint+tls+certIssuer, die man einzeln raten kann und dann eine Seite hat, die halb funktioniert). Endet nicht bei "deployed", sondern bei URL, Benutzer, Passwort und DNS-Record. Die behaltenen Volumes sind richtig so — eine Datenbank an einen Vertipper zu verlieren wäre schlimmer. Nur ihre Folge stand nirgends: sie machen die nächste "frische Installation" zu einem Upgrade auf altem Bestand, ohne neues Passwort. Das Skript erkennt sie und fragt; Löschen verlangt ein getipptes "delete", --yes deckt es bewusst nicht ab. Zwei Guards in `make check`, beide vor dem Tag statt danach: check-changelog.sh — der CI-Guard, der 0.1.70 versenkt hat check-commands.sh — kein "…" in einem Shell-Block, den jemand ausführen soll 0.1.70 wird nicht nachgezogen: der Tag ist öffentlich, veröffentlicht wurde darunter nichts, und eine Lücke ist ein ehrlicheres Protokoll als ein verschobener Tag. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXXVNHRBuPjPJbo6L5eg3z
1 parent 4d56fc8 commit 4e421a6

10 files changed

Lines changed: 978 additions & 14 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 52 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -15,6 +15,58 @@ matching image, so a version identifies one exact pair
1515
1616
## [Unreleased]
1717

18+
## [0.1.71] — 2026-09-05
19+
20+
### Added
21+
22+
- **An installer**: `scripts/install.sh` — install, upgrade, uninstall, and read the
23+
admin password back. It checks the things that actually go wrong before Helm runs:
24+
a kubeconfig that is not on `$KUBECONFIG` (k3s writes a root-only one), an
25+
unreachable API server, a missing ingress controller, and data left behind by a
26+
previous install. Then it asks how TLS should be terminated — cert-manager,
27+
Cloudflare *Full*, Cloudflare *Flexible*, or plain HTTP — and afterwards prints the
28+
URL, the user, the password and the DNS record needed to reach it.
29+
Run it with `bash <(curl -fsSL …/scripts/install.sh)`; prompts are read from
30+
`/dev/tty`, so piping the script does not make it answer its own questions.
31+
- **`--delete-data`**, and a matching prompt in `uninstall`, for the volumes
32+
`helm uninstall` deliberately keeps. Keeping them is right — losing a database to a
33+
typo is worse — but it silently turns the next "fresh install" into an upgrade over
34+
an old database, which is how an install can complete with no admin password to show
35+
for it. Deleting them takes the flag or a typed confirmation; `--yes` does not
36+
cover it.
37+
38+
### Fixed
39+
40+
- **Restoring an archive failed on the schema's own foreign keys.** The first test to
41+
take a real archive from a real database and put it back found three defects in ten
42+
seconds, after six releases of green unit tests: rows were inserted children-first,
43+
each table was truncated separately (which Postgres refuses for a referenced table
44+
even when the child is empty), and blank CSV fields became SQL `NULL` in `NOT NULL`
45+
columns. Insert order and truncation now come from the target's `pg_constraint`
46+
graph, and nullability from `information_schema`, so a new foreign key needs no one
47+
to remember this.
48+
- **A fresh install could finish with no way to log in.** The bootstrap password was
49+
printed to the pod log exactly once, and `secrets.adminPassword` was only read when
50+
the admin user was created — so it did nothing in the one case it was needed. It is
51+
now generated into the release Secret like the database password and the JWT key,
52+
kept across upgrades, and applied on **every** start, which makes setting it the way
53+
to reset it. Read it with `scripts/install.sh password`.
54+
- **The config seed defaulted to a path from the developer's machine**
55+
(`/root/ess-config-values`), so every fresh install logged a `permission denied`
56+
warning about a directory that was never theirs. The default is now empty.
57+
- **A README command could not be pasted**: it contained a literal `…` where the rest
58+
of the command belonged, and pasting it produced
59+
`non-absolute URLs should be in form of repo_name/path_to_chart, got: %E2%80%A6`.
60+
61+
### Notes
62+
63+
- **0.1.70 was tagged but never published.** The release workflow stopped at its
64+
"CHANGELOG must have a section for this version" guard — correctly, but after the
65+
tag was already public, so nothing reached GHCR and `helm install` kept resolving
66+
0.1.69 while everything local reported success. The same check now runs in
67+
`make check` (`scripts/check-changelog.sh`), before a tag exists. The version number
68+
is skipped rather than reused: a gap is a more honest record than a moved tag.
69+
1870
## [0.1.69] — 2026-09-05
1971

2072
### Fixed

‎Makefile‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -56,6 +56,8 @@ check:
5656
$(GO) test ./...
5757
cd web && ./node_modules/.bin/tsc -b --noEmit
5858
./scripts/check-sensitive.sh
59+
./scripts/check-changelog.sh
60+
./scripts/check-commands.sh
5961
# gofmt is a CI gate, and `make check` did not run it until 2026-08-17 — so
6062
# "check green" did not imply "CI green", and E51 shipped unformatted code that
6163
# only the pipeline would have caught. A local check that omits a remote gate

‎README.md‎

Lines changed: 63 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -141,16 +141,51 @@ helm version
141141
**One more thing if you want HTTPS.** The install command below passes
142142
`ingress.certIssuer=letsencrypt-prod`, which assumes [cert-manager](https://cert-manager.io/docs/installation/)
143143
and a `ClusterIssuer` of that name already exist. Install cert-manager first, or
144-
drop the `--set ingress.certIssuer=…` flag and terminate TLS however you prefer.
144+
drop the `--set ingress.certIssuer=<issuer>` flag and terminate TLS however you prefer.
145145

146146
Both installer scripts above are piped straight from the internet into a shell.
147147
That is what the upstream projects document, but read them first if that is not
148148
acceptable in your environment.
149149
</details>
150150

151-
### Install (recommended) — OCI chart
151+
### Install (recommended) — the installer script
152152

153-
The chart and image are published to GHCR, so one command is all you need:
153+
```bash
154+
bash <(curl -fsSL https://raw.githubusercontent.com/bxnnyg/matrixctrl/master/scripts/install.sh)
155+
```
156+
157+
It asks for a hostname and how you want HTTPS terminated, and checks the rest itself:
158+
whether `kubectl`/`helm` are there, whether a kubeconfig exists where nothing looks for
159+
it (k3s writes a root-only one to `/etc/rancher/k3s/k3s.yaml`), whether the cluster
160+
answers, whether an ingress controller is installed, whether cert-manager has a
161+
`ClusterIssuer` — and whether a previous install left a database behind, which is the
162+
one thing that quietly turns a reinstall into an upgrade.
163+
164+
When it finishes it prints the URL, the user, the **admin password**, and the DNS
165+
record you need. That is the part a bare `helm install` leaves you to find.
166+
167+
Non-interactive, for a script or a second server:
168+
169+
```bash
170+
./scripts/install.sh install --host matrixctrl.example.com \
171+
--tls letsencrypt --issuer letsencrypt-prod --yes
172+
```
173+
174+
Other subcommands: `password` (prints it again, any time), `status`, `uninstall`.
175+
176+
#### HTTPS — pick one
177+
178+
| `--tls` | What happens | When |
179+
|---|---|---|
180+
| `letsencrypt` | cert-manager issues a real certificate | cert-manager and a `ClusterIssuer` exist, and port 80 is reachable. **Not behind Cloudflare's proxy** — HTTP-01 cannot complete through it |
181+
| `cloudflare-full` | Traefik serves its own default certificate; Cloudflare re-encrypts | Cloudflare proxied, SSL mode **Full**. Nothing to install. Not *Full (strict)* |
182+
| `cloudflare-flexible` | The origin speaks plain HTTP | Cloudflare proxied, SSL mode **Flexible** |
183+
| `none` | Plain HTTP | LAN, or a tunnel that terminates TLS itself |
184+
185+
<details>
186+
<summary><b>Alternative — plain Helm, no script</b></summary>
187+
188+
The chart and image are published to GHCR:
154189

155190
```bash
156191
helm install matrixctrl oci://ghcr.io/bxnnyg/charts/matrixctrl \
@@ -159,6 +194,11 @@ helm install matrixctrl oci://ghcr.io/bxnnyg/charts/matrixctrl \
159194
--set ingress.certIssuer=letsencrypt-prod
160195
```
161196

197+
On k3s, `export KUBECONFIG=/etc/rancher/k3s/k3s.yaml` first — without it Helm reports
198+
`Kubernetes cluster unreachable: dial tcp [::1]:8080` and means "I did not find your
199+
cluster".
200+
</details>
201+
162202
No version is pinned here on purpose: Helm resolves the newest published chart, so
163203
this command cannot go stale. Each released chart pins its own matching image, so
164204
"newest chart" still means one exact, reproducible pair — not a moving `latest`.
@@ -212,7 +252,7 @@ helm install matrixctrl oci://ghcr.io/bxnnyg/charts/matrixctrl \
212252
To choose it yourself — at install **or any time after** — set it and upgrade:
213253

214254
```bash
215-
helm upgrade matrixctrl … --set secrets.adminPassword='your-password'
255+
./scripts/install.sh install --admin-password 'your-password'
216256
```
217257

218258
The value is applied on every start, so this is also how you reset a password you
@@ -286,13 +326,13 @@ helm upgrade matrixctrl oci://ghcr.io/bxnnyg/charts/matrixctrl -n matrixctrl --r
286326
--set secrets.adminPassword='your-password'
287327
```
288328

289-
> **Upgrading to v0.1.70 sets a new bootstrap password.** The chart generates one into
329+
> **Upgrading to v0.1.71 sets a new bootstrap password.** The chart generates one into
290330
> its Secret because there was none there before, and the app applies it on start — so
291331
> read it back with the command above rather than reusing the old one. This affects only
292332
> the local `admin` bootstrap login; if you have already switched to **Connect Matrix
293333
> Login**, that is not how you sign in and nothing changes for you.
294334
295-
> Before v0.1.70 the password was written to the pod log exactly once and
335+
> Before v0.1.71 the password was written to the pod log exactly once and
296336
> `secrets.adminPassword` was read only when the account was created — so a restarted pod
297337
> meant no way in, and reinstalling did not help because `helm uninstall` keeps the
298338
> database volume. If you are on an older version, upgrade; that alone fixes it.
@@ -321,19 +361,32 @@ scale it back up.
321361
### Uninstall
322362

323363
```bash
324-
helm uninstall matrixctrl -n matrixctrl
364+
./scripts/install.sh uninstall
325365
```
326366

327-
This deliberately **leaves three things behind**, so that reinstalling does not
328-
lose your data or lock you out:
367+
It removes the release, then shows you what Helm kept and offers to delete that
368+
too — typing `delete` is required, and `--yes` does not cover it.
369+
370+
By hand it is `helm uninstall matrixctrl -n matrixctrl`, which deliberately
371+
**leaves three things behind**, so that reinstalling does not lose your data or
372+
lock you out:
329373

330374
| Kept | Why |
331375
|---|---|
332376
| `pvc/matrixctrl-config` | the git config repo — every version and rollback point |
333377
| `pvc/matrixctrl-postgres` | audit log, hooks, upgrade history |
334378
| `secret/matrixctrl-secret` | DB password and JWT key — regenerating them invalidates every session |
335379

336-
They carry `helm.sh/resource-policy: keep`. To remove everything for real:
380+
They carry `helm.sh/resource-policy: keep`.
381+
382+
> **This is why "just reinstall it" does not reset anything.** With those volumes
383+
> in place a fresh `helm install` comes up on the existing database: the schema is
384+
> already migrated, the admin account already exists, and the install has no new
385+
> password to give you — it looks like a first install and behaves like an upgrade.
386+
> If that is what you are trying to escape, delete the volumes below, or use
387+
> `./scripts/install.sh install`, which notices them and asks.
388+
389+
To remove everything for real:
337390

338391
```bash
339392
kubectl delete pvc matrixctrl-config matrixctrl-postgres -n matrixctrl

‎deploy/helm/matrixctrl/Chart.yaml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2,8 +2,8 @@ apiVersion: v2
22
name: matrixctrl
33
description: Admin layer for self-hosted Matrix / Element Server Suite (ESS)
44
type: application
5-
version: 0.1.70
6-
appVersion: "0.1.70"
5+
version: 0.1.71
6+
appVersion: "0.1.71"
77
keywords:
88
- matrix
99
- element

‎docs/DESIGN.md‎

Lines changed: 80 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2780,3 +2780,83 @@ The general shape, for the next time: **a secret that exists in exactly one plac
27802780
exactly one moment, is not stored — it is announced.** Anything an operator will need
27812781
later has to live somewhere they can query on their own schedule, not somewhere they had
27822782
to be watching.
2783+
2784+
### §4.75 — Der Fix war fertig, getestet und kam nie an (2026-09-05, operator + agent, etappe 76)
2785+
2786+
Letzte Runde habe ich §4.74 gemeldet als *ausgeliefert*: `make check` grün, beide
2787+
Integrationstests grün, Image gebaut, in k3s importiert, `helm upgrade` auf Revision 73,
2788+
Pod 2/2, Logzeile aus dem laufenden Container zitiert. Jede einzelne dieser Aussagen war
2789+
wahr.
2790+
2791+
Der Operator hat auf einer frischen VM installiert und bekam `0.1.69`.
2792+
2793+
**Der Release-Job war an Schritt 6 gescheitert:** *"CHANGELOG must have a section for
2794+
this version"*. Ich hatte getaggt, ohne den Abschnitt zu schreiben. Der Guard hat exakt
2795+
das getan, wofür er gebaut wurde — nur steht er hinter dem Tag. Ein Tag ist gepusht und
2796+
öffentlich, bevor irgendetwas ihn prüft; was danach fehlschlägt, ist ein Bericht, kein
2797+
Schutz. GHCR blieb auf 0.1.69, `helm install` löste brav „die neueste veröffentlichte
2798+
Version" auf, und der Operator lief in genau den Bug, den 0.1.70 behebt.
2799+
2800+
Nachgesehen habe ich es nicht. `gh` fehlte auf der Maschine, und dabei habe ich es
2801+
belassen — statt des einen `curl` auf `api.github.com/repos/…/actions/runs`, der zwei
2802+
Sekunden dauert und die Frage beantwortet hätte. „Verifiziert" hieß hier: alles geprüft,
2803+
was ich lokal anfassen konnte, und die eine Sache nicht, die zwischen meiner Maschine und
2804+
seiner steht.
2805+
2806+
Die Verwechslung dahinter ist benennbar: **in den Cluster deployt ist nicht
2807+
veröffentlicht.** Der lokale k3s ist die Maschine, auf der ich arbeite — dass dort etwas
2808+
läuft, sagt über die Registry nichts. Alle Belege, die ich gesammelt hatte, kamen von der
2809+
falschen Seite der Veröffentlichungsgrenze.
2810+
2811+
Zwei Konsequenzen, beide vor dem Tag statt danach:
2812+
2813+
- `scripts/check-changelog.sh` in `make check` — derselbe Guard wie in CI, an der Stelle,
2814+
wo er noch etwas verhindern kann statt es zu protokollieren.
2815+
- Eine Release-Meldung gilt erst, wenn die Registry sie bestätigt. Der Tag-Push ist die
2816+
Absicht, nicht das Ergebnis.
2817+
2818+
0.1.70 wird nicht nachgezogen. Der Tag ist öffentlich, veröffentlicht wurde unter der
2819+
Nummer nichts, und eine Lücke in der Versionsfolge ist ein ehrlicheres Protokoll als ein
2820+
verschobener Tag — die Nummern 0.1.36, 0.1.42 und 0.1.63 fehlen aus demselben Grund.
2821+
2822+
### §4.76 — Vier Fehlschläge auf dem Weg zu einer Installation, keiner davon in der Software (2026-09-05, operator, etappe 76)
2823+
2824+
Das Terminal des Operators, in seiner Reihenfolge:
2825+
2826+
| | Was er sah | Was es hieß |
2827+
|---|---|---|
2828+
| 1 | `Kubernetes cluster unreachable: Get "http://localhost:8080/version"` | `KUBECONFIG` nicht gesetzt. k3s legt sie root-only unter `/etc/rancher/k3s/k3s.yaml` ab |
2829+
| 2 | `These resources were kept due to the resource policy` | `helm uninstall` behält die Volumes. Die Neuinstallation lief auf der alten Datenbank |
2830+
| 3 | Passwort-Secret leer | 0.1.70 war nicht veröffentlicht (§4.75) |
2831+
| 4 | `non-absolute URLs …, got: %E2%80%A6` | Er hat ein README-Kommando mit einem echten `…` darin eingefügt |
2832+
2833+
Vier Fehlschläge, keiner in MatrixCtrl. Alle vier im Weg dorthin — und jeder einzelne war
2834+
dokumentiert. Die Kubeconfig-Zeile steht im README. Was `helm uninstall` behält, steht im
2835+
README, mit Tabelle und Begründung. Trotzdem hat er zwei Stunden verloren, weil
2836+
Dokumentation zum Zeitpunkt des Fehlers nur hilft, wenn man weiß, wonach man sucht: eine
2837+
Fehlermeldung über Port 8080 schickt niemanden in einen Abschnitt über Kubeconfig-Pfade.
2838+
2839+
Was fehlte, war nichts Erklärendes, sondern etwas Ausführbares. `scripts/install.sh`
2840+
prüft der Reihe nach genau das, was hier schiefging, und sagt bei jedem Punkt einen Satz
2841+
statt eines Exit-Codes. Die interessanteste Stelle ist Nr. 2: die behaltenen Volumes sind
2842+
**richtig** so — eine Datenbank an einen Vertipper zu verlieren wäre schlimmer —, aber
2843+
ihre Folge wurde nie ausgesprochen. Sie machen die nächste „frische Installation" zu
2844+
einem Upgrade auf altem Bestand: Schema migriert, Admin-Account vorhanden, kein neues
2845+
Passwort zu zeigen. Genau der Zustand, aus dem der Operator herauswollte, war der, den
2846+
Neuinstallieren herstellt. Das Skript fragt jetzt, statt zu raten, und verlangt für das
2847+
Löschen ein getipptes `delete` — `--yes` deckt es bewusst nicht ab: ein Flag, das „hör
2848+
auf zu fragen" bedeutet, darf keine Datenbank löschen.
2849+
2850+
Nr. 4 hat einen eigenen Guard bekommen (`scripts/check-commands.sh`): in einem
2851+
Shell-Block darf kein `…` stehen. Prosa darf abkürzen; ein Block, den jemand ausführen
2852+
soll, ist eine Anweisung, und eine Anweisung mit einer Lücke ist eine Falle für genau den
2853+
Leser, der nicht weiß, was in die Lücke gehört.
2854+
2855+
Die TLS-Frage stellt das Skript aus, weil der Operator sie gestellt hat: cert-manager
2856+
(HTTP-01 scheitert hinter Cloudflares Proxy), Cloudflare *Full* (Traefiks
2857+
Default-Zertifikat reicht, nichts zu installieren), Cloudflare *Flexible* (Origin spricht
2858+
HTTP), oder gar kein TLS. Vier Kombinationen aus `entrypoint`/`tls`/`certIssuer`, die man
2859+
einzeln raten kann und dann eine Seite hat, die halb funktioniert.
2860+
2861+
Der Lauf endet nicht bei `STATUS: deployed`, sondern bei URL, Benutzer, Passwort und dem
2862+
DNS-Record. Eine Installation ist fertig, wenn jemand eingeloggt ist.

‎docs/ROADMAP.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -60,8 +60,9 @@ Etappes 1–10 are **reconstructed from `git log`** (39 commits, 2026-05-27 →
6060
| 32 | Release Notes auf der Upgrade-Seite + Version aus der Liste übernommen — die andere Hälfte der Pin-Warnung | ✅ 2026-08-05 · `v0.1.33` · [plan](plans/etappe-32-release-notes.md) |
6161
| 33 | OIDC-Init wiederholen statt einmalig aufgeben — ein Neustart vor MAS sperrte den Operator 11 h aus dem eigenen Panel aus | ✅ 2026-08-06 · `v0.1.34` · [plan](plans/etappe-33-oidc-retry.md) |
6262
| 73 | Der Restore, der nichts wiederhergestellt und Erfolg gemeldet hätte — dazu Skeletons und ein ehrlicher Byte-Zähler | ✅ 2026-09-05 · `v0.1.69` · [plan](plans/etappe-73-silent-restore-and-loading.md) |
63-
| 74 | Der Restore, der an den eigenen Fremdschlüsseln scheiterte — gefunden vom ersten echten Durchlauf | ✅ 2026-09-05 · `v0.1.70` · [plan](plans/etappe-74-restore-order.md) |
64-
| 75 | Ein frisches Setup, in das man sich auch anmelden kann — Passwort im Secret statt in einer Logzeile | ✅ 2026-09-05 · `v0.1.70` · [plan](plans/etappe-75-first-login.md) |
63+
| 74 | Der Restore, der an den eigenen Fremdschlüsseln scheiterte — gefunden vom ersten echten Durchlauf | ✅ 2026-09-05 · `v0.1.71` · [plan](plans/etappe-74-restore-order.md) |
64+
| 75 | Ein frisches Setup, in das man sich auch anmelden kann — Passwort im Secret statt in einer Logzeile | ✅ 2026-09-05 · `v0.1.71` · [plan](plans/etappe-75-first-login.md) |
65+
| 76 | Ein Installer, der die vier Fallen auf dem Weg kennt — und ein Release, das ankommt statt nur zu existieren | ✅ 2026-09-05 · `v0.1.71` · [plan](plans/etappe-76-installer.md) |
6566
| 72 | Ein Backup statt drei Entschuldigungen | ✅ 2026-09-05 · `v0.1.68` · [plan](plans/etappe-72-one-backup.md) |
6667
| 71 | Wo die Konfiguration wirklich liegt — die beruhigendste Eigenschaft, die nie jemand ausgesprochen hat | ✅ 2026-09-05 · `v0.1.67` · [plan](plans/etappe-71-where-the-config-lives.md) |
6768
| 70 | Die Volumes — und ein Navigationspunkt, der ins Leere zeigte | ✅ 2026-09-05 · `v0.1.66` · [plan](plans/etappe-70-homeserver-export.md) |

0 commit comments

Comments
 (0)