You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Keep usage event records of running apps, service instances, and tasks
The scheduled usage event cleanup job used to delete every record older than
the configured cutoff age, including the opening STARTED/CREATED event of a
resource that is still running. Once the cleanup deleted that event, nothing
was left to reconstruct what is running right now.
Database::OldRecordCleanup can now optionally keep the records of running
resources. Each model declares its lifecycles via usage_lifecycles: which
states open a run (STARTED/CREATED/TASK_STARTED, plus the
WAS_RUNNING/TASK_WAS_RUNNING baselines), which state closes it
(STOPPED/DELETED/TASK_STOPPED), and which column names the resource. An old
opening event is then only deleted when:
* a closing event for the same resource exists later and is also old -- the
run is over; or
* it is neither the first opening of the current run nor the resource's
latest one (again judged only against old rows). Consumers only need the
first opening (the true start time) and the latest (the current size). The
ones in between, written each time a running resource is scaled or updated,
tell a consumer nothing it still needs -- and deleting them is what keeps
the table size bounded for long-running, frequently-changed resources.
The app and service usage event repositories turn this on with
keep_running_records: true. Asking for it on a model without usage_lifecycles
raises an error instead of silently deleting the records of running
resources. Task events get their own lifecycle (TASK_STARTED/TASK_WAS_RUNNING
-> TASK_STOPPED, matched by task_guid), so the start events of long-running
tasks survive cleanup too. Task baselines use their own TASK_WAS_RUNNING
state because task events carry an empty app_guid: if they said WAS_RUNNING,
the app lifecycle would see them all as events of one app whose guid is ''
and wrongly delete them (and the backfill's repair would write bogus STOPPED
events for that phantom app).
Deletion runs in a deliberate order: first the opening events that are safe
to delete, while the events that make them safe still exist; then everything
else. The reverse order could delete a closing event first and leave its
opening event looking like a still-running resource. The cleanup log line now
reports the row counts BatchDelete returns instead of running extra COUNT
queries, and BatchDelete fetches each batch's ids in the same query that
checks whether anything is left, halving the evaluations of the (potentially
expensive) filtered dataset. Also renames the positional days_ago argument to
a cutoff_age_in_days keyword.
0 commit comments