Skip to content

Commit a7fd1c9

Browse files
committed
Backfill WAS_RUNNING events for running apps, tasks, and service instances
Seed a synthetic WAS_RUNNING usage event for every currently-running app process, a TASK_WAS_RUNNING event for every currently-running task, and a WAS_RUNNING event for every existing service instance. Billing consumers can then bootstrap a complete picture of what is running, even though the usage event cleanup deleted the original STARTED/TASK_STARTED/CREATED events long ago. The backfill is a batched VCAP::WasRunningBackfill helper called from thin no_transaction migrations, following the bigint-migration pattern. It walks the started processes / running tasks / service instances in id order, one batch at a time, each batch in its own READ COMMITTED transaction -- so no statement comes near the migration statement timeout, and MySQL's INSERT..SELECT takes no shared next-key locks on the scanned rows while the API keeps serving traffic. Tasks in CANCELING count as running: they stay billable until Diego reports them dead, and no usage event marks the moment a task enters CANCELING. The app query limits its package/droplet subqueries to each batch's apps so it never scans those whole tables, and it COALESCEs nullable legacy columns so one bad NULL row cannot abort a deploy. The seeds skip any resource whose start is already on record -- an earlier baseline, or a real STARTED/TASK_STARTED/CREATED/UPDATED event -- so running the backfill again cannot give a resource a second start that a consumer would bill twice. The API stays live during migrations, so a seed batch can race a stop or delete and write a baseline for a resource that is already gone -- or whose stop event landed earlier in the table, with a lower id. Deleting such rows would not help: consumers read these tables forward, by id, and keep what they read. A poller may already have the baseline, and for tasks a TASK_STOPPED may already have been written against it. You can delete a row; you cannot make a consumer un-read it. So instead, a post-seed repair adds the missing ending event (STOPPED / DELETED / TASK_STOPPED) for every baseline whose resource is no longer running and that has no later ending event (one with a higher id). The ending is built from the baseline row itself, which carries every NOT NULL column an ending needs -- necessary, because the resource row may be gone entirely. A baseline that already has its real ending is never touched, and each added ending stops its baseline from matching the test, so re-running the backfill changes nothing. Two properties of the added ending are deliberate. Its created_at is the repair time, not the true stop time: a bounded overbill that ends, which beats a missing ending billed forever. And its previous_state is the baseline's state, which no normal ending carries, so repaired endings are easy to tell apart. A skip_was_running_backfill config flag lets operators opt out. The migrations check it (not the helper), because they are recorded as applied either way; 'rake db:was_running_backfill' runs the same seeding and repair later. Use the rake task after a skipped migration, once after the deploy that ships these migrations (to repair anything that slipped through while old API servers were still running), or after a destructive usage-event purge, which wipes the task start events that task stop events depend on. The post-deploy run matters because the seed migrations run at the start of a rolling deploy, while old API servers still serve traffic and their old cleanup code can still delete start events the new code depends on. The last seed migration logs this reminder, so it shows up in the migration output operators see during the deploy. The rake task rejects a batch size that is not a positive whole number, instead of silently seeding nothing. Every backfill run holds a session advisory lock, so two runs cannot both add the same missing ending. The seed migrations wait for the lock: a deploy that pauses behind a finishing rake run is harmless, and a failed deploy is not. The rake task fails fast instead, so an operator gets feedback rather than a silent queue. This is the first advisory lock in this codebase; a comment in the helper explains why the locking tools already in use do not fit this job. The migrations' down blocks are deliberate no-ops: consumers may already have read the seeded rows, and deleting a row cannot make a consumer un-read it -- it would only leave the stop events written against these rows without a start event to pair with. Document the WAS_RUNNING/TASK_WAS_RUNNING states, their created_at semantics, the repaired ending events, and the rules consumers must follow on the V3 resources, and list the new states in the legacy V2 usage-event docs because V2 reads the same event rows.
1 parent e30e8af commit a7fd1c9

21 files changed

Lines changed: 1796 additions & 2 deletions
Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,27 @@
1+
require 'database/was_running_backfill'
2+
3+
Sequel.migration do
4+
no_transaction # backfill manages its own per-batch transactions
5+
6+
up do
7+
logger = Steno.logger('cc.backfill.was_running')
8+
if VCAP::WasRunningBackfill.skip?
9+
VCAP::WasRunningBackfill.log_skip(logger, 'app')
10+
else
11+
# wait: an operator's 'rake db:was_running_backfill' may be running; a
12+
# deploy that briefly pauses behind it is harmless, failing it is not.
13+
VCAP::WasRunningBackfill.with_advisory_lock(self, wait: true) do
14+
VCAP::WasRunningBackfill.seed_app_usage_events(self, logger)
15+
end
16+
end
17+
end
18+
19+
down do
20+
# Deliberately a no-op. Consumers may already have read the seeded rows,
21+
# and deleting a row cannot make a consumer un-read it -- it would only
22+
# leave any later STOPPED events without a start event to pair with.
23+
# Leaving the rows is safe: re-running the migration or the
24+
# 'db:was_running_backfill' rake task skips resources that already have a
25+
# baseline.
26+
end
27+
end
Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,27 @@
1+
require 'database/was_running_backfill'
2+
3+
Sequel.migration do
4+
no_transaction # backfill manages its own per-batch transactions
5+
6+
up do
7+
logger = Steno.logger('cc.backfill.was_running')
8+
if VCAP::WasRunningBackfill.skip?
9+
VCAP::WasRunningBackfill.log_skip(logger, 'service')
10+
else
11+
# wait: an operator's 'rake db:was_running_backfill' may be running; a
12+
# deploy that briefly pauses behind it is harmless, failing it is not.
13+
VCAP::WasRunningBackfill.with_advisory_lock(self, wait: true) do
14+
VCAP::WasRunningBackfill.seed_service_usage_events(self, logger)
15+
end
16+
end
17+
end
18+
19+
down do
20+
# Deliberately a no-op. Consumers may already have read the seeded rows,
21+
# and deleting a row cannot make a consumer un-read it -- it would only
22+
# leave any later DELETED events without a start event to pair with.
23+
# Leaving the rows is safe: re-running the migration or the
24+
# 'db:was_running_backfill' rake task skips instances that already have a
25+
# baseline.
26+
end
27+
end
Lines changed: 38 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,38 @@
1+
require 'database/was_running_backfill'
2+
3+
Sequel.migration do
4+
no_transaction # backfill manages its own per-batch transactions
5+
6+
up do
7+
logger = Steno.logger('cc.backfill.was_running')
8+
if VCAP::WasRunningBackfill.skip?
9+
VCAP::WasRunningBackfill.log_skip(logger, 'task')
10+
else
11+
# wait: an operator's 'rake db:was_running_backfill' may be running; a
12+
# deploy that briefly pauses behind it is harmless, failing it is not.
13+
VCAP::WasRunningBackfill.with_advisory_lock(self, wait: true) do
14+
VCAP::WasRunningBackfill.seed_task_usage_events(self, logger)
15+
end
16+
# This is the last of the three seed migrations, so remind the operator
17+
# here. The migrations run at the start of a rolling deploy, while old
18+
# API servers are still serving traffic; their old cleanup code can
19+
# still delete start events that the new code depends on. The seed
20+
# cannot see the future, so the operator has to close that window by
21+
# running the backfill once more after the deploy finishes.
22+
logger.info("WAS_RUNNING usage event backfill complete. If old API servers were still serving traffic during this deploy, run 'rake db:was_running_backfill' " \
23+
'once after the deploy finishes to repair anything they changed in the meantime.')
24+
end
25+
end
26+
27+
down do
28+
# Deliberately a no-op. Consumers may already have read the seeded rows,
29+
# and deleting a row cannot make a consumer un-read it -- it would only
30+
# leave any later TASK_STOPPED events without a start event to pair with.
31+
# Worse: a task's stop event is only written when the task has recorded
32+
# start evidence, and these rows ARE that evidence for tasks whose
33+
# TASK_STARTED the cleanup already deleted. Remove them and those tasks'
34+
# eventual stops are silently swallowed. Leaving the rows is safe:
35+
# re-running the migration or the 'db:was_running_backfill' rake task
36+
# skips tasks that already have a baseline.
37+
end
38+
end

‎docs/v2/app_usage_events/list_all_app_usage_events.html‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -631,9 +631,11 @@ <h4>Body</h4>
631631
<ul class="valid_values">
632632
<li>STARTED</li>
633633
<li>STOPPED</li>
634+
<li>WAS_RUNNING</li>
634635
<li>BUILDPACK_SET</li>
635636
<li>TASK_STARTED</li>
636637
<li>TASK_STOPPED</li>
638+
<li>TASK_WAS_RUNNING</li>
637639
</ul>
638640
</td>
639641
<td>

‎docs/v2/service_usage_events/list_service_usage_events.html‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -290,6 +290,7 @@ <h4>Body</h4>
290290
<li>CREATED</li>
291291
<li>DELETED</li>
292292
<li>UPDATED</li>
293+
<li>WAS_RUNNING</li>
293294
</ul>
294295
</td>
295296
<td>

‎docs/v3/source/includes/resources/app_usage_events/_delete.md.erb‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,8 @@ Content-Type: application/json
2121

2222
Destroys all existing events. Populates new usage events, one for each started app. All populated events will have a `created_at` value of current time. There is the potential race condition if apps are currently being started, stopped, or scaled. The seeded usage events will have the same guid as the app.
2323

24+
**Note:** the reseed only writes `STARTED` events for app processes — it does not restore the start evidence (`TASK_STARTED`/`TASK_WAS_RUNNING`) of currently-running tasks, and `TASK_STOPPED` events are only emitted for tasks with recorded start evidence. After a purge, operators should run `rake db:was_running_backfill` on a Cloud Controller VM to reseed baselines for running tasks; otherwise their eventual stops are silently suppressed.
25+
2426
#### Definition
2527
`POST /v3/app_usage_events/actions/destructively_purge_all_and_reseed`
2628

‎docs/v3/source/includes/resources/app_usage_events/_object.md.erb‎

Lines changed: 27 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -30,3 +30,30 @@ Name | Type | Description
3030
**instance_count.current** | _integer_ or `null` | Current instance count of the app that this event pertains to, if applicable
3131
**instance_count.previous** | _integer_ or `null` | Previous instance count of the app that this event pertains to, if applicable
3232
**links** | [_links object_](#links) | Links to related resources
33+
34+
#### WAS_RUNNING and TASK_WAS_RUNNING events
35+
36+
`WAS_RUNNING` and `TASK_WAS_RUNNING` are synthetic values for `state.current` recorded once per running process (`WAS_RUNNING`) and once per running task (`TASK_WAS_RUNNING`). They are written by the upgrade migration that introduced the keep-running cleanup feature, and again whenever an operator runs the `db:was_running_backfill` task (for example after a purge); a resource that already has a start event on record is skipped. They mark every process and task that was already running at the time of the upgrade so that billing consumers can bootstrap from a complete baseline even if the original `STARTED`/`TASK_STARTED` events have been pruned.
37+
38+
**Consumer interpretation** (read `WAS_RUNNING`/`STARTED` as `TASK_WAS_RUNNING`/`TASK_STARTED` for task events, which are keyed by `task.guid`):
39+
40+
* If you have not previously recorded a `STARTED` event for this resource, treat `WAS_RUNNING` as equivalent to `STARTED`.
41+
* If you have already recorded `STARTED` (or an earlier `WAS_RUNNING`) for this resource, treat as a redundant baseline confirmation and ignore.
42+
* `created_at` reflects when the backfill migration ran, **not** when the app or task actually started. Treat `WAS_RUNNING` as a baseline marker that the resource was already running as of that timestamp, not as the true start of the running interval.
43+
* `state.previous` on a `WAS_RUNNING` event is always `null`. Subsequent real events for the same resource will continue to report their actual prior process state in `state.previous` (typically `STARTED`). If you perform chain validation, treat `WAS_RUNNING` as equivalent to `STARTED` for the purpose of validating the next event's `state.previous`.
44+
45+
#### Repaired ending events
46+
47+
The backfill (and any later run of its recovery task) repairs baselines that turn out to be unpaired: if a `WAS_RUNNING`/`TASK_WAS_RUNNING` event was recorded for a resource that is no longer running and no later ending event exists for it — for example because the resource stopped while the backfill was still in progress — the missing ending event (`STOPPED`/`TASK_STOPPED`) is appended. Baselines are never deleted. A repaired ending event:
48+
49+
* carries a `created_at` of when the repair ran, **not** when the resource actually stopped — the interval it closes may overstate the true run by that gap;
50+
* copies the footprint (`instance_count`, `memory_in_mb_per_instance`) of the baseline it pairs;
51+
* reports the baseline's state (`WAS_RUNNING`/`TASK_WAS_RUNNING`) in `state.previous`, which normal ending events never carry — use this to tell repaired endings apart.
52+
53+
#### What a consumer must do
54+
55+
Independent of the backfill, the events stream asks three things of any consumer that pairs beginnings with endings:
56+
57+
1. Ignore a `WAS_RUNNING`/`TASK_WAS_RUNNING` event for a resource you already track (see above).
58+
2. Tolerate duplicate ending events for the same resource: close the interval on the first ending after a beginning and ignore further endings until the next beginning.
59+
3. Treat an ending event with no visible beginning for that resource as noise.

‎docs/v3/source/includes/resources/service_usage_events/_object.md.erb‎

Lines changed: 22 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -26,3 +26,25 @@ Name | Type | Description
2626
**service_broker.guid** | _string_ or `null` | Unique identifier of the service broker that this event pertains to, if applicable
2727
**service_broker.name** | _string_ or `null` | Name of the service broker that this event pertains to, if applicable
2828
**links** | [_links object_](#links) | Links to related resources
29+
30+
#### WAS_RUNNING events
31+
32+
`WAS_RUNNING` is a synthetic value for `state` recorded once per existing service instance. It is written by the upgrade migration that introduced the keep-running cleanup feature, and again whenever an operator runs the `db:was_running_backfill` task (for example after a purge); an instance that already has a `CREATED`, `UPDATED`, or `WAS_RUNNING` event on record is skipped. It marks every service instance that existed at the time of the upgrade so that billing consumers can bootstrap from a complete baseline of service instances even if the original `CREATED` events have been pruned.
33+
34+
**Consumer interpretation:**
35+
36+
* If you have not previously recorded a `CREATED` event for this service instance, treat `WAS_RUNNING` as equivalent to `CREATED`.
37+
* If you have already recorded `CREATED` (or an earlier `WAS_RUNNING`) for this instance, treat as a redundant baseline confirmation and ignore.
38+
* `created_at` reflects when the backfill migration ran, **not** when the service instance was created. Treat `WAS_RUNNING` as a baseline marker that the instance already existed as of that timestamp.
39+
40+
#### Repaired ending events
41+
42+
The backfill (and any later run of its recovery task) repairs baselines that turn out to be unpaired: if a `WAS_RUNNING` event was recorded for a service instance that no longer exists and no later `DELETED` event exists for it — for example because the instance was deleted while the backfill was still in progress — the missing `DELETED` event is appended, copying the baseline's instance, plan, and broker attributes. Baselines are never deleted. A repaired `DELETED` event carries a `created_at` of when the repair ran, **not** when the instance was actually deleted — the interval it closes may overstate the instance's true lifetime by that gap.
43+
44+
#### What a consumer must do
45+
46+
Independent of the backfill, the events stream asks three things of any consumer that pairs beginnings with endings:
47+
48+
1. Ignore a `WAS_RUNNING` event for a service instance you already track (see above).
49+
2. Tolerate duplicate `DELETED` events for the same instance: close the interval on the first one and ignore the rest.
50+
3. Treat a `DELETED` event with no visible `CREATED`/`UPDATED`/`WAS_RUNNING` for that instance as noise.

‎lib/cloud_controller/config_schemas/api_schema.rb‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -109,6 +109,7 @@ class ApiSchema < VCAP::Config
109109
optional(:migration_psql_concurrent_statement_timeout_in_seconds) => Integer,
110110
optional(:migration_psql_worker_memory_kb) => Integer,
111111
optional(:skip_bigint_id_migration) => bool,
112+
optional(:skip_was_running_backfill) => bool,
112113
db: {
113114
optional(:database) => Hash, # db connection hash for sequel
114115
max_connections: Integer, # max connections in the connection pool

‎lib/cloud_controller/config_schemas/migrate_schema.rb‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ class MigrateSchema < VCAP::Config
1010
optional(:migration_psql_concurrent_statement_timeout_in_seconds) => Integer,
1111
optional(:migration_psql_worker_memory_kb) => Integer,
1212
optional(:skip_bigint_id_migration) => bool,
13+
optional(:skip_was_running_backfill) => bool,
1314

1415
db: {
1516
optional(:database) => Hash, # db connection hash for sequel

0 commit comments

Comments
 (0)