Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
83 changes: 65 additions & 18 deletions plugins/tend-ci-runner/scripts/list-recent-runs.sh
Original file line number Diff line number Diff line change
Expand Up @@ -97,23 +97,50 @@ if [ -n "$cron_minute" ]; then
# Default floor: one cron period back. Consecutive ticks tile exactly.
COMPLETED_AFTER=$((intended - 3600))

# Dropped-tick recovery. GHA doesn't only *delay* scheduled ticks, it also
# *drops* them: a tick that fires zero times leaves that hour's completions
# in the gap between the previous and next cycle's windows (the skipped-tick
# case #526 deferred as acceptable). Rather than assume the previous tick
# fired, resume from where the previous *actual* completed run of this
# workflow left off: recover that run's intended tick and floor the window
# there. When every tick fires, the previous run's intended tick == the
# Dropped-tick and failed-run recovery. GHA doesn't only *delay* scheduled
# ticks, it also *drops* them: a tick that fires zero times leaves that hour's
# completions in the gap between the previous and next cycle's windows (the
# skipped-tick case #526 deferred as acceptable). Rather than assume the
# previous tick fired, resume from where the previous run that actually
# analyzed something left off: recover that run's intended tick and floor the
# window there.
#
# Anchor on the previous *successful* run, not merely a completed one. A run
# that failed before its agent produced any analysis — harness crash, model
# quota exhaustion, timeout — covered no window at all, so resuming from it
# discards every hour the outage spanned and the next green run reports a
# one-hour all-clear for a window nothing ever looked at. `--status` takes
# conclusions as well as statuses, so the filter runs server-side and no scan
# depth is involved: an outage of any length can't bury the anchor.
#
# Scheduled runs only. A `workflow_dispatch` run has no `.schedule` in its
# event payload, so it took the non-cron branch and covered a now-anchored 1h
# window rather than a tiled one — anchoring on it would discard the rest of
# the gap, silently, which is the same class of hole as anchoring on a
# failure. Dispatching the workflow by hand to check on a fix mid-outage is
# the natural thing to do, and is exactly when that would bite. A
# partially-failed matrix run counts as a failure here, which only ever widens
# the window — overlap is re-offered work the caller dedups against its own
# evidence log, where a gap is silently unanalyzed.
#
# When every tick fires and succeeds, the previous run's intended tick == the
# default (intended - 3600), so this is a byte-identical no-op — still no
# overlap between consecutive cycles. When a tick was dropped, it reaches
# back to cover the orphaned hour. Capped at 6h so a sustained outage can't
# create an unbounded window. The analyzing workflow runs on the current
# repo, so this query omits TARGET_REPO's -R.
# overlap between consecutive cycles. Capped at 6h so a sustained outage can't
# create an unbounded window; when the cap bites, say so on stderr so the
# caller records a coverage gap instead of a false all-clear. The analyzing
# workflow runs on the current repo, so this query omits TARGET_REPO's -R.
if [ -n "${GITHUB_WORKFLOW:-}" ]; then
prev_start=$(gh run list --workflow "$GITHUB_WORKFLOW" --status completed \
--limit 10 --json databaseId,createdAt \
--jq "[.[] | select(.databaseId != (${GITHUB_RUN_ID:-0}))] | .[0].createdAt // empty" \
2>/dev/null || true)
floor_cap=$((intended - 21600)) # never reach back more than 6h
floor_cap_iso=$(date -u -d "@$floor_cap" +%Y-%m-%dT%H:%M:%SZ)
# Route through gh_retry and fail loud, like the other queries here: a
# transient API error must not masquerade as "no successful run ever", which
# would emit a confidently-worded warning naming a cause that didn't happen.
if ! prev_start=$(gh_retry gh run list --workflow "$GITHUB_WORKFLOW" \
--status success --event schedule --limit 5 --json databaseId,createdAt \
--jq "[.[] | select(.databaseId != (${GITHUB_RUN_ID:-0}))] | .[0].createdAt // empty"); then
echo "ERROR: 'gh run list' for the window anchor failed after retries — refusing to guess a floor that would read as a false all-clear" >&2
exit 1
fi
if [ -n "$prev_start" ]; then
prev_ts=$(date -u -d "$prev_start" +%s 2>/dev/null || echo "")
if [ -n "$prev_ts" ]; then
Expand All @@ -123,10 +150,15 @@ if [ -n "$cron_minute" ]; then
else
prev_intended=$((prev_hour_tick - 3600))
fi
floor_cap=$((intended - 21600)) # never reach back more than 6h
[ "$prev_intended" -lt "$floor_cap" ] && prev_intended=$floor_cap
if [ "$prev_intended" -lt "$floor_cap" ]; then
echo "WARNING: the last successful '$GITHUB_WORKFLOW' run started $prev_start, more than 6h back. Window floored at $floor_cap_iso; runs that completed before it are NOT in this list. Record a coverage gap, not an all-clear." >&2
prev_intended=$floor_cap
fi
[ "$prev_intended" -lt "$COMPLETED_AFTER" ] && COMPLETED_AFTER=$prev_intended
fi
else
echo "WARNING: no successful scheduled '$GITHUB_WORKFLOW' run found at all. Window floored at $floor_cap_iso; anything earlier is NOT in this list. Record a coverage gap, not an all-clear." >&2
[ "$floor_cap" -lt "$COMPLETED_AFTER" ] && COMPLETED_AFTER=$floor_cap
fi
fi

Expand All @@ -138,16 +170,31 @@ fi

all_runs="[]"

# `gh run list` returns newest-first, so a workflow with more runs in the window
# than this limit silently drops the *oldest* ones — exactly the runs sitting in
# a gap the recovery above just reached back for, which would restore the false
# all-clear that recovery exists to prevent. A recovered window spans up to 8h,
# and a single busy workflow clears 50 runs in that span, so the limit is sized
# well past one workflow's output; gh paginates in hundreds, so the headroom
# costs at most one extra page per workflow.
RUN_LIMIT=200

for wf in "${WORKFLOWS[@]}"; do
if ! runs=$(gh_retry gh run list \
"${repo_args[@]}" \
--workflow "${wf}" \
--created ">=${CREATED_SINCE}" \
--json databaseId,conclusion,createdAt,updatedAt \
--limit 50); then
--limit "$RUN_LIMIT"); then
echo "ERROR: 'gh run list' for workflow '$wf' failed after retries — refusing to report a partial run list" >&2
exit 1
fi
# Exactly $RUN_LIMIT rows means the list may be capped rather than complete.
# Warn rather than exit: unlike a failed fetch, the rows in hand are still
# worth analyzing — the caller just can't read a short list as an all-clear.
if [ "$(printf '%s' "$runs" | jq 'length')" -ge "$RUN_LIMIT" ]; then
echo "WARNING: '$wf' returned $RUN_LIMIT runs, the fetch limit — older runs in this window are likely missing from the list. Record a coverage gap, not an all-clear." >&2
fi
all_runs=$(echo "$all_runs" "$runs" | jq -s 'add')
done

Expand Down
2 changes: 2 additions & 0 deletions plugins/tend-ci-runner/skills/review-reviewers/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -215,6 +215,8 @@ The script discovers `tend-*` workflows by default. Pass additional prefixes as

If empty, record the run as all-clear per "Recording below-threshold findings" above, then skip to Step 6.

If the script printed a `WARNING:` on stderr, the list is known-incomplete — the window was clamped, no anchor was found, or a workflow hit the fetch limit. Record a coverage gap naming the missing span instead of an all-clear, whether or not the list came back empty; the next run's floor advances past that span regardless, so an unrecorded gap is never revisited.

## Step 2: Survey outcomes via cheap subagent

Spawn a cheap subagent to check outcomes across all runs from Step 1. The subagent does the token-heavy work of mapping runs to PRs/issues and checking acceptance signals.
Expand Down
Loading