-
Notifications
You must be signed in to change notification settings - Fork 1
500 lines (467 loc) · 27.2 KB
/
Copy pathbot-serving-check.yml
File metadata and controls
500 lines (467 loc) · 27.2 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
# Synthetic monitor for the nginx bot -> seo-proxy path.
#
# Humans get the SPA shell and never notice when the bot hop breaks: the site
# looks healthy, Plausible shows normal traffic, CI is green — and every
# crawler sees an error page. Exactly that happened between 2026-06-12 and
# 2026-07-09: the @seo_proxy upstream TLS verification failed (default
# proxy_ssl_verify_depth 1 vs a 4-deep Let's Encrypt chain) and every bot UA
# received HTTP 502 on every page, undetected for ~4 weeks. This check is the
# alarm that was missing.
#
# The checks target the Cloud Run ORIGIN, not https://anyplot.ai: Cloudflare's
# bot management 403s GitHub-runner (datacenter) IPs — including UA-spoofed
# "Googlebot", which only passes verified-bot checks from real Google IPs —
# verified on the first dispatched run. The origin is also exactly the layer
# that broke in the incident above; Cloudflare-edge issues are out of this
# monitor's reach by design.
#
# Skipping the edge is exactly what the app's origin gate refuses
# (app/origin-gate.conf.template), so every probe below carries the header the
# edge would have stamped, out of the ORIGIN_SECRET repository secret. This
# monitor is the reason the gate could not be a Host rule: it cannot spoof the
# Host either, and anything else it might present as an exception would be
# public with this repository.
#
# What the bot pages are is DERIVED from the repo, never written out here. The
# routes come from the `@router.get("/seo-proxy/…")` decorators in
# api/routers/seo.py and the spec title from plots/<spec>/specification.yaml.
# Hand-written literals go stale: the home title changed to "anyplot.ai —
# AI-generated plot catalog for 15 libraries" and this check sat red for ten
# consecutive nights on a healthy site — an alarm nobody can trust, which is
# worse than no alarm. Deriving turns the check into what it is meant to assert
# ("what is served == what the repo says"), covers every bot page the moment it
# lands in seo.py, and cannot drift on a copy change again.
name: Bot Serving Check
on:
schedule:
- cron: "23 6 * * *" # daily 06:23 UTC
workflow_dispatch:
# One run at a time, and the newest wins. Two overlapping runs — a manual
# dispatch during the nightly one, which can take up to the timeout below —
# would both mutate the same alarm issue: two failures could open it twice, and
# an older run finishing last could close a newer failure or reopen an alarm
# after a newer success. Cancelling the superseded run means only the latest
# result ever touches the issue.
concurrency:
group: bot-serving-check-${{ github.ref }}
cancel-in-progress: true
# issues: write is the alarm path. Without it the job could only go red in a
# tab nobody subscribes to — which is how those ten red nights stayed unnoticed.
# The failure step opens (or comments on) one fixed-title issue and the success
# step closes it again.
permissions:
contents: read
issues: write
jobs:
bot-serving:
runs-on: ubuntu-latest
# 36 check() calls (10 derived bot routes + spec page + impl page +
# ClaudeBot + 15 crawler UAs + 404 + robots + sitemap + 3 llms + 2 human
# controls) x (--retry 2 -> up to 3 attempts x --max-time 30) can reach
# ~54 min worst-case, plus the retried /_health gate probe (90s) and five
# non-retried probes (llms.txt charset,
# trailing slash, og-image, .well-known redirect, the /{spec}/{language}
# 301 — 30s each); 64 leaves
# room to report a clean failure rather than dying to the job timeout,
# which reports nothing useful. Recompute this when adding checks: the
# ceiling is check() calls x 90s, plus margin. The route sweep grows with
# api/routers/seo.py, so a new bot page adds 90s to that ceiling.
timeout-minutes: 64
steps:
# The routes and the expected title are read out of the repo, so it has
# to be here before the first request goes out.
- name: Check out the repo
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 1
- name: Crawler UAs must get 200 + per-route pages
env:
# The header the Cloudflare Transform Rule stamps on everything it
# proxies for anyplot.ai. These checks target the Cloud Run ORIGIN on
# purpose (see the file header: Cloudflare 403s runner IPs even for a
# UA-spoofed Googlebot), so they skip the edge and have to stamp their
# own — the same repository secret sync-postgres.yml sends to the API
# gate, and the same value both services are given. It reaches curl
# through the environment and this step runs without `set -x`, so it
# is never echoed. An exception keyed on anything else — a header this
# workflow invents, a user agent — would be public with this
# repository, which is why the shared secret is the only exception the
# gate can afford.
ORIGIN_SECRET: ${{ secrets.ORIGIN_SECRET }}
run: |
set -uo pipefail
# Cloud Run origin of the anyplot-app service (see header comment
# for why not https://anyplot.ai).
ORIGIN="https://anyplot-app-r3tvmejsmq-ez.a.run.app"
GOOGLEBOT="Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)"
TWITTERBOT="Twitterbot/1.0"
CLAUDEBOT="Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)"
CHATGPTUSER="Mozilla/5.0 (compatible; ChatGPT-User/1.0; +https://openai.com/bot)"
HUMAN="Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0 Safari/537.36"
fail=0
# Absent secret => no header at all. Sending an empty one instead
# would read as a mismatch, i.e. as a half-applied rotation rather
# than as the missing secret it is.
HDR=()
if [ -n "${ORIGIN_SECRET}" ]; then HDR=(-H "X-Origin-Secret: ${ORIGIN_SECRET}"); fi
# Ask the gate before anything else. /_health is its one exempt path,
# so it answers whatever the gate is doing, and it reports the verdict
# for the request it was asked with — never the value. Without this,
# an armed gate plus a missing repository secret would red all ~36
# checks below with 403s and open an incident that reads "every
# crawler page is broken" — the wrong thing to be woken up for, and
# the wrong thing to go looking for.
gate=$(curl -sS --retry 2 --max-time 30 -o /dev/null -D - "${HDR[@]}" "$ORIGIN/_health" \
| tr -d '\r' | sed -n 's/^[Xx]-[Oo]rigin-[Gg]ate: //p') || gate=""
case "$gate" in
ok|off-seen)
echo "origin gate: $gate — the probes below carry the header" ;;
off)
echo "::warning::the origin gate is not armed AND this run sent no X-Origin-Secret. Set the ORIGIN_SECRET repository secret to the value the anyplot-app service will be given, before it is armed." ;;
missing)
echo "::error::the origin gate is armed and this run has no ORIGIN_SECRET repository secret — every check below would be a 403. Set it to the same value as the ORIGIN_SECRET on the anyplot-app Cloud Run service."
exit 1 ;;
mismatch)
echo "::error::the origin gate is armed and the ORIGIN_SECRET repository secret is not the value the anyplot-app service was given — a half-applied rotation. Roll the gate back, finish the rotation, arm again."
exit 1 ;;
*)
# No header at all: the running revision predates the gate. That
# is deploy lag, not a serving fault, and every check below still
# means exactly what it always meant — so it is said out loud and
# the run continues.
echo "::warning::$ORIGIN/_health carries no X-Origin-Gate header, so the deployed revision predates the origin gate. Check for a pending deploy." ;;
esac
check() {
local ua="$1" url="$2" expect="$3" want="${4:-200}"
local code
# On curl failure REPLACE the code — a failing curl can still have
# printed a partial -w code; appending would yield e.g. "200000".
code=$(curl -sS --retry 2 --max-time 30 "${HDR[@]}" -A "$ua" -o body.html -w '%{http_code}' "$url") || code="000"
if [ "$code" != "$want" ]; then
echo "::error::$url with UA '$ua' returned HTTP $code (expected $want)"
fail=1
elif ! grep -qF "$expect" body.html; then
echo "::error::$url with UA '$ua' returned $want but the body is missing: $expect"
fail=1
else
echo "OK: $url ($ua)"
fi
}
# Bot path: prerendered per-route HTML from the seo-proxy.
#
# The routes are the API's own bot pages — every
# `@router.get("/seo-proxy/…")` in api/routers/seo.py is a page a
# crawler must be able to reach. Reading them from the source means a
# page added there is swept the moment it lands, and no list here can
# go stale.
#
# The extraction parses the AST rather than grepping for a literal
# spelling: a decorator written with single quotes, wrapped over two
# lines, or carrying kwargs would be silently skipped by a regex, and
# a silently short list is exactly the failure this file is trying to
# stop making. The parameterised routes (paths containing `{`) cannot
# be swept generically and are each probed by hand below, so their
# COUNT is returned too and checked against what this file probes —
# a fourth one added to seo.py fails the run instead of going
# unnoticed.
#
# The assertion is the canonical link, not a title: it is GENERATED
# from the route, so it can never drift on a copy change, and it
# proves both halves at once — the SPA shell carries no
# `<link rel="canonical">` at all, so a match means the bot hop ran,
# and the href names the route, so it means the right page came back.
seo_decorators=$(python3 - <<'PY'
import ast
SRC = "api/routers/seo.py"
static, templated, unresolved = set(), 0, 0
for node in ast.walk(ast.parse(open(SRC, encoding="utf-8").read())):
for dec in getattr(node, "decorator_list", []):
if not isinstance(dec, ast.Call):
continue
fn = dec.func
if not (isinstance(fn, ast.Attribute) and fn.attr == "get"):
continue
if not (isinstance(fn.value, ast.Name) and fn.value.id == "router"):
continue
# FastAPI takes the path positionally OR as `path=`. Anything
# else — a variable, an f-string, a constant folded elsewhere —
# is a route this cannot reason about, and staying silent about
# it is how a bot page would drop out of the sweep unnoticed.
arg = dec.args[0] if dec.args else next(
(kw.value for kw in dec.keywords if kw.arg == "path"), None
)
if not isinstance(arg, ast.Constant) or not isinstance(arg.value, str):
unresolved += 1
continue
path = arg.value
if not path.startswith("/seo-proxy"):
continue
if "{" in path:
templated += 1
else:
static.add(path.removeprefix("/seo-proxy") or "/")
for route in sorted(static):
print("route " + route)
print("templated %d" % templated)
print("unresolved %d" % unresolved)
PY
)
routes=$(printf '%s\n' "$seo_decorators" | sed -n 's/^route //p')
templated=$(printf '%s\n' "$seo_decorators" | sed -n 's/^templated //p')
unresolved=$(printf '%s\n' "$seo_decorators" | sed -n 's/^unresolved //p')
swept=0
while IFS= read -r route; do
[ -n "$route" ] || continue
swept=$((swept + 1))
check "$GOOGLEBOT" "$ORIGIN$route" "<link rel=\"canonical\" href=\"https://anyplot.ai$route\" />"
done <<< "$routes"
# A checkout that silently produced nothing, a parse error, or a
# refactor that moves the decorators somewhere this cannot see would
# otherwise pass the whole sweep by checking zero routes.
if [ "$swept" -lt 1 ]; then
echo "::error::no static bot routes found in api/routers/seo.py — the sweep asserted nothing"
fail=1
else
echo "swept $swept derived bot route(s)"
fi
# The parameterised routes are probed one by one below: /{spec_id},
# /{spec_id}/{language} and /{spec_id}/{language}/{library}. Bump this
# number only together with a new probe for the new route.
if [ "$templated" != "3" ]; then
echo "::error::api/routers/seo.py has $templated parameterised /seo-proxy route(s); this monitor probes 3 by hand — add a probe for the new one and update this guard"
fail=1
fi
# A router.get whose path is not a string literal cannot be classified
# at all, and a route this cannot see is a route it cannot promise to
# cover. Loud, not silent.
if [ "$unresolved" != "0" ]; then
echo "::error::$unresolved router.get decorator(s) in api/routers/seo.py carry a path this monitor cannot resolve — it may be skipping a bot page"
fail=1
fi
# The spec page's expected title comes from the spec file, not from a
# literal here. `<title>$SPEC_TITLE` as a prefix: the suffix is site
# copy ("| anyplot.ai") and the impl page inserts the library name,
# neither of which this monitor is the right place to pin.
# Two decodings, both mirroring what the page actually does with the
# value. yaml.safe_load, not the source line: a quoted title is valid
# YAML and already occurs in this repo (plots/heatmap-chromagram).
# html.escape, because api/routers/seo.py escapes the title before it
# reaches `<title>` — `Mohr's Circle` is served as `Mohr's
# Circle`, and a raw needle would be a false alarm against a perfectly
# healthy page. Both are the same failure this rewrite exists to
# remove: an expectation that does not match what is served.
SPEC="scatter-basic"
python3 -c "import yaml" 2>/dev/null || pip install --quiet --disable-pip-version-check pyyaml
SPEC_TITLE=$(python3 - "$SPEC" <<'PY'
import html
import sys
import yaml
with open("plots/%s/specification.yaml" % sys.argv[1], encoding="utf-8") as fh:
doc = yaml.safe_load(fh) or {}
title = doc.get("title")
print(html.escape(title) if isinstance(title, str) else "")
PY
)
if [ -z "$SPEC_TITLE" ]; then
echo "::error::plots/$SPEC/specification.yaml carries no title — nothing to assert against"
fail=1
SPEC_TITLE="__no_title_in_the_spec_file__"
else
echo "expecting the $SPEC pages to be titled: $SPEC_TITLE"
fi
check "$GOOGLEBOT" "$ORIGIN/$SPEC" "<title>$SPEC_TITLE"
# The implementation page is asserted on its canonical instead: a
# three-segment route is the one that would still look right with the
# hub page served in its place, and the title prefix cannot tell them
# apart.
check "$TWITTERBOT" "$ORIGIN/$SPEC/python/matplotlib" \
"<link rel=\"canonical\" href=\"https://anyplot.ai/$SPEC/python/matplotlib\" />"
# The middle tier, /{spec}/{language}, was consolidated onto the hub
# and must answer 301 -> /{spec}. Its own docstring records why it is
# worth a probe: a Location of /seo-proxy/{spec} is re-prefixed by
# nginx, arrives back at this route and redirects forever — Googlebot
# logged 48 "Redirect error" URLs before the target was sanitised. The
# loop signature is a target that still carries /seo-proxy.
# The STATUS is part of the contract, not only the target: a 302/307/308
# to the same hub would consolidate nothing, and the endpoint documents
# a permanent redirect.
read -r lang_code lang_target <<< "$(curl -sS --max-time 30 -o /dev/null \
"${HDR[@]}" -A "$GOOGLEBOT" -w '%{http_code} %{redirect_url}' "$ORIGIN/$SPEC/python")"
if [ "$lang_code" != "301" ]; then
echo "::error::/$SPEC/python answered HTTP $lang_code (expected a permanent 301 onto the hub)"
fail=1
else
case "$lang_target" in
*"/seo-proxy"*)
echo "::error::/$SPEC/python redirects into the internal proxy path: $lang_target"
fail=1 ;;
"$ORIGIN/$SPEC") echo "OK: /$SPEC/python -> 301 $lang_target" ;;
*)
echo "::error::/$SPEC/python should 301 to $ORIGIN/$SPEC, got '$lang_target'"
fail=1 ;;
esac
fi
# AI assistants take the same prerendered path (nginx $is_bot). These
# checks hit the ORIGIN, so they verify the nginx map independently of
# whether Cloudflare's AI Crawl Control currently 403s these UAs at
# the edge — an edge-level policy change needs no change here.
check "$CLAUDEBOT" "$ORIGIN/$SPEC" "<title>$SPEC_TITLE"
# User-directed fetchers: a human asked their assistant to open the
# page. All of these were verified receiving the empty SPA shell on
# 2026-08-18 — an assistant asked about a plot could describe nothing.
# Google documents its own as generally ignoring robots.txt, so the
# nginx map is the only control point and the only thing this guards.
for ua in \
"Mozilla/5.0 (compatible; Google-GeminiNotebook)" \
"Mozilla/5.0 (compatible; Google-NotebookLM)" \
"Mozilla/5.0 (compatible; Gemini-Deep-Research)" \
"Mozilla/5.0 (compatible; GoogleAgent-Mariner)" \
"meta-externalfetcher/1.1" \
"Mozilla/5.0 (compatible; Meta-WebIndexer/1.0)" \
"Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)" \
"Mozilla/5.0 (compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)" \
"DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)" \
"Mozilla/5.0 (compatible; Amzn-SearchBot/1.0)" \
"Mozilla/5.0 (compatible; Amzn-User/1.0)" \
"Mozilla/5.0 (compatible; Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)" \
"meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)" \
"Grok/1.0" \
"Mozilla/5.0 (compatible; xAI-Bot/1.0)"
do
check "$ua" "$ORIGIN/$SPEC" "<title>$SPEC_TITLE"
done
# A crawler asking for a URL that is no page gets a real 404 from the
# seo-proxy — the SPA shell would answer 200 (soft-404), and did for
# 161 stale migration URLs before the proxy learned to say no.
check "$GOOGLEBOT" "$ORIGIN/this-spec-does-not-exist" '"status":404' 404
# The machine files must be served directly, never proxied to the
# seo backend — including for a mapped crawler UA, which is the whole
# point of the `location =` bypasses. robots.txt and sitemap.xml
# join llms.txt below: the sitemap is proxied to the API for every
# client, and a broken proxy would hand a crawler the SPA shell.
check "$GOOGLEBOT" "$ORIGIN/robots.txt" "User-agent: Bytespider"
check "$GOOGLEBOT" "$ORIGIN/sitemap.xml" "<urlset"
# The site card must be the FILE for a preview bot, not the proxy —
# a preview bot that lands on /seo-proxy/og-image.png shows nothing.
code=$(curl -sS --max-time 30 "${HDR[@]}" -A "$TWITTERBOT" -o /dev/null -w '%{http_code} %{content_type}' "$ORIGIN/og-image.png") || code="000"
case "$code" in
"200 image/png"*) echo "OK: og-image.png served as image to a preview bot" ;;
*) echo "::error::og-image.png for a preview bot: $code (expected 200 image/png)"; fail=1 ;;
esac
# A guessed /.well-known/llms.txt must land on the file, not on the
# SPA shell (which soft-404'd it with 200 until 2026-08-28).
wk_target=$(curl -sS --max-time 30 -o /dev/null "${HDR[@]}" -A "$CHATGPTUSER" \
-w '%{redirect_url}' "$ORIGIN/.well-known/llms.txt")
case "$wk_target" in
*/llms.txt) echo "OK: .well-known/llms.txt -> $wk_target" ;;
*) echo "::error::.well-known/llms.txt did not redirect to the guide: '$wk_target'"; fail=1 ;;
esac
# llms.txt must be served directly, never proxied to the seo backend —
# including for a mapped crawler UA, which is the whole point of the
# `location = /llms.txt` bypass.
check "$GOOGLEBOT" "$ORIGIN/llms.txt" "# anyplot"
check "$CHATGPTUSER" "$ORIGIN/llms.txt" "# anyplot"
# llms-full.txt is proxied to the API for EVERY client (mapped or
# not) — before, the SPA catch-all soft-404'd it with the homepage
# shell. The catalogue-index line proves the API generated it.
check "$HUMAN" "$ORIGIN/llms-full.txt" "# anyplot — full catalogue index"
# llms.txt carries UTF-8 punctuation (em dashes, arrows); without an
# explicit charset a strict client decodes it as Latin-1 mojibake.
ct=$(curl -sS --max-time 30 "${HDR[@]}" -A "$GOOGLEBOT" -o /dev/null -w '%{content_type}' "$ORIGIN/llms.txt")
case "$ct" in
*charset=utf-8*) echo "OK: llms.txt content-type: $ct" ;;
*)
echo "::error::llms.txt served without utf-8 charset: $ct"
fail=1 ;;
esac
# A trailing slash must normalise to the canonical URL on THIS host.
# It used to 307 to http://api.anyplot.ai/seo-proxy/... — internal
# path, wrong host, plain http, and that host disallows all crawling.
slash_target=$(curl -sS --max-time 30 -o /dev/null "${HDR[@]}" -A "$GOOGLEBOT" \
-w '%{redirect_url}' "$ORIGIN/$SPEC/")
case "$slash_target" in
"")
# No redirect at all: %{redirect_url} is empty, which the previous
# form printed as "OK: trailing slash -> " and passed. Copilot
# raised this twice; it means the rewrite has disappeared.
echo "::error::trailing slash produced no redirect — the rewrite is gone"
fail=1 ;;
*"/seo-proxy"*|http://*|*:8080/*)
echo "::error::trailing-slash redirect leaks or downgrades: $slash_target"
fail=1 ;;
*) echo "OK: trailing slash -> $slash_target" ;;
esac
# Control: humans must still get the SPA shell — on the home page
# and on a deep route.
check "$HUMAN" "$ORIGIN/" '<div id="root">'
check "$HUMAN" "$ORIGIN/$SPEC" '<div id="root">'
exit $fail
# The alarm. One issue with a fixed title carries the whole history of
# this monitor: a first failure opens it, every further failure comments
# on it (so a long outage is one thread, not one issue per night), and
# the first green run closes it again.
#
# Only from the default branch. The schedule always runs there, but
# workflow_dispatch can pick any branch that carries this file — and a
# branch experimenting with the derived expectations must not be able to
# raise, or silently close, a repository-wide production incident. Off
# main the checks still run and still red the job; only the issue is left
# alone.
- name: Raise the alarm
if: failure() && github.ref == format('refs/heads/{0}', github.event.repository.default_branch)
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
ALARM_TITLE: Bot serving check is red
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
run: |
set -euo pipefail
# env.ALARM_TITLE inside the jq filter, not string interpolation:
# the title travels as data and never as jq syntax.
num=$(gh issue list --state open --limit 100 --json number,title \
--jq '[.[] | select(.title == env.ALARM_TITLE)] | .[0].number // empty')
if [ -n "$num" ]; then
echo "alarm issue #$num is already open — appending this run"
gh issue comment "$num" --body "Still red: $RUN_URL"
exit 0
fi
# printf over one argument per line: a YAML block scalar cannot carry
# unindented lines, and a leading run of spaces inside the body would
# render as a markdown code block on the issue.
body=$(printf '%s\n' \
"The daily bot-serving monitor failed: $RUN_URL" \
"" \
"Crawlers may be getting the SPA shell or an error page while the site" \
"looks healthy to humans — the failure mode that went unnoticed for four" \
"weeks in 2026 (see the header of the workflow file)." \
"" \
"Two causes are worth separating before anything else:" \
"" \
'- **The serving path broke** — `app/nginx.conf` (the `$is_bot` map, the' \
' `@seo_proxy` upstream) or the API `/seo-proxy` routes. A real incident.' \
'- **The deployed pages are behind the repo** — the swept routes come from' \
' `api/routers/seo.py` and the expected title from the spec file, so a' \
' merged change that has not been deployed yet reads as a mismatch. Check' \
' for a pending deploy first.' \
"" \
"This issue closes itself on the next green run.")
echo "opening the alarm issue"
gh issue create --title "$ALARM_TITLE" --label bug --label infrastructure --body "$body"
- name: Stand the alarm down
if: success() && github.ref == format('refs/heads/{0}', github.event.repository.default_branch)
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
GH_REPO: ${{ github.repository }}
ALARM_TITLE: Bot serving check is red
RUN_URL: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}
run: |
set -euo pipefail
num=$(gh issue list --state open --limit 100 --json number,title \
--jq '[.[] | select(.title == env.ALARM_TITLE)] | .[0].number // empty')
if [ -n "$num" ]; then
echo "closing alarm issue #$num"
gh issue comment "$num" --body "Green again: $RUN_URL"
gh issue close "$num"
else
echo "no open alarm issue — nothing to close"
fi