Skip to content

Concepts Alerts

Arael Espinosa edited this page Jul 18, 2026 · 2 revisions

Alerts

An Alert is the runtime record of something that fired — either a threshold rule (Alert Config) firing on a specific Check, or a notification received from a third-party system via an inbound Integration (currently: GCP Cloud Monitoring). It is a distinct entity from an Incident — an Alert tracks "this condition has been active since X"; an Incident is a manually declared, user-facing disruption. Alerts never create or attach to Incidents automatically — that link is always a deliberate action taken from the Alerts panel.

A check can have multiple Alert Configs. Since RFC 0002, Alert Configs are N:1 with Checks — one check can have several severity rules attached to it (for example, a "Warning" rule and a separate "Critical" rule for the same metric). See Alert Configs below.

Key properties

Field Type Description
AlertConfigId int? The threshold rule that fired. Null for an External alert — there is no internal Alert Config behind a third-party occurrence
CheckId int? The check whose result triggered this alert. Null for an External alert
ServiceId int? The service the check belongs to. Null for an External alert
IncidentId int? Set once this alert is linked to an incident (null until that manual step happens)
ImpactAtFireTime ServiceStatus The impact snapshot at the moment the alert fired — derived from the Alert Config's severity (internal) or mapped from the source's own severity (external)
Message string? Optional detail captured at fire time
MessageFingerprint string Dedup key for internal alerts — see Occurrence tracking
FiredAt datetime When the alert first fired
ResolvedAt datetime? When the alert resolved (recovery threshold met, or a closed event for an External alert). Null while still active
OccurrenceCount int How many times this same alert condition has re-fired without resolving — see below
EscalationCurrentStep / EscalationStepStartedAt int?, datetime? Tracks progress through the assigned Escalation Policy, if any
EscalationPolicyId int? The escalation policy snapshotted for this alert at creation time — see Escalation snapshot
AcknowledgedAt / AcknowledgedBy long?, string? Set when a team member acknowledges the alert. Blocked once the alert is Resolved — see Lifecycle
LastUserActivityAt datetime? Last time a human interacted with this alert — feeds escalation re-triggers
Source enum Where the alert came from — Internal (Piro's own check evaluation) or GcpCloudMonitoring. See External alerts
ExternalId string? The source system's own identifier for this occurrence (e.g. GCP's incident_id). Null for Internal alerts, which dedup via MessageFingerprint instead
SourceUrl string? Optional deep link back into the source system's own console for this occurrence (e.g. GCP Cloud Monitoring's incident URL). Shown as a "View in GCP Cloud Monitoring"-style link on the alert detail page when present
SourceRequestLogId int? Points at the exact inbound webhook request that produced this alert, if any — see WebhookRequestLog

Source: Alert.cs

External alerts

Most alerts are anchored to a Check and a Service, as described above. An alert can also arrive with no Check and no Service at all — the admin UI calls this an External alert (an alert "managed externally of Piro"). This happens today only via the GCP Cloud Monitoring webhook integration: Piro receives a notification about something GCP's own alerting policy decided to fire, with nothing in Piro to correlate it against.

For an External alert:

  • AlertConfigId, CheckId, and ServiceId are all null.
  • Fields that only make sense for a threshold rule — AlertFor, AlertValue, and severity thresholds — are not applicable. The admin UI shows "Managed externally of Piro" in the Trigger Criteria section instead of those fields.
  • Dedup uses (Source, ExternalId) instead of MessageFingerprint — repeated deliveries for the same source-side incident update the same Alert rather than creating duplicates.
  • It can still escalate to on-call — see Escalation snapshot below.
  • It can still be promoted to an Incident, and the admin explicitly chooses which Service(s), if any, the resulting Incident should be attached to — see Manual linking to an Incident.

🔍 Correlating an incoming GCP alert to a specific Piro Check/Service is not implemented yet — every alert from the GCP Cloud Monitoring integration is currently External. This is a deliberate scope decision for the first version of that integration, not a limitation of the Alert model itself.

Escalation snapshot

EscalationPolicyId on an Alert is set once, at creation time, from whichever policy applies:

  • For a normal (Check-anchored) alert: the Service's escalationPolicyId.
  • For an External alert: the source Integration's own EscalationPolicyId, if the admin configured one on that integration.

This is a frozen snapshot, not a live reference — editing the policy afterward does not retroactively change the behavior of an alert that's already partway through escalation.

Alert Configs

An Alert Config is the threshold rule attached to a check. As of RFC 0002, checks no longer judge their own severity — a check's executor only reports whether it measured successfully (UP) or failed to measure at all (DOWN, e.g. connection refused, DNS didn't resolve, TLS handshake failed), plus a raw metric value (latency in ms, days until certificate expiry, count of failed name servers, etc.). Deciding what that raw value means — "this counts as degraded," "this counts as critical," and at what threshold — is entirely the Alert Config's job.

Alert Configs are N:1 with a Check

A single check can have any number of Alert Configs, each independently evaluating the check's raw metric. This lets you layer severities on the same signal instead of picking one threshold — for example, on an SSL check:

Alert Config AlertFor AlertValue Severity
"Warning — renew soon" CertExpiry 30 (days) Warning
"Critical — expiring imminently" CertExpiry 7 (days) Critical

Both configs watch the same MetricValue (days remaining) reported by the check; each fires independently once the certificate crosses its own threshold. Previously (pre-RFC 0002) a check could only have one Alert Config — this is the key behavior change to be aware of when migrating existing checks.

Key properties

Field Type Description
CheckId int The check this rule is attached to
AlertFor enum The metric this rule evaluates — see below
AlertValue string The threshold: a ServiceStatus name for Status alerts, milliseconds for Latency, days for CertExpiry, or a failed-name-server count for FailedNameServers
Severity enum Warning or Critical
FailureThreshold int Consecutive failing measurements before the alert fires (default 1)
SuccessThreshold int Consecutive passing measurements before the alert auto-resolves (default 1)
IsActive bool Disables the rule without deleting it
Description string? Optional free-text note

Source: AlertConfig.cs

AlertFor types

AlertFor Compares against Meaning
Status the check's UP/DOWN result Fires based on the executor's own success/failure determination
Latency CheckDataPoint.LatencyMs Fires when response/resolution time crosses AlertValue (ms)
CertExpiry CheckDataPoint.MetricValue Fires when days-until-expiry drops to or below AlertValue. SSL checks only
FailedNameServers CheckDataPoint.MetricValue Fires when the count of name servers that failed to resolve reaches AlertValue. DNS checks only

🔍 The old Uptime AlertFor was removed — it was never actually implemented and has no replacement; use Status for availability-based alerting.

Source: AlertFor.cs, AlertSeverity.cs

Which AlertFor values are valid per check type

Not every AlertFor makes sense for every Check type — a GCP Cloud Run Job check has no latency signal, and only SSL checks report certificate expiry. The admin UI only offers the values valid for the check's type, enforced server-side:

CheckType Allowed AlertFor values
HTTP Status, Latency
DNS Status, Latency, FailedNameServers
TCP Status, Latency
Ping Status, Latency
SSL Status, CertExpiry
Heartbeat Status
GRPC Status, Latency
GCP_CloudRunJob Status

Source: CheckTypeExtensions.cs

No Alert Config is created automatically

Creating a check does not create a default Alert Config. If you save a check without adding at least one Alert Config, it will run on schedule and record data, but nothing will ever notify you — there is no built-in fallback severity rule.

⚠️ The admin panel warns you about this. When creating a check with zero Alert Configs configured, the check-creation form shows a confirmation dialog before letting you proceed, so this is a deliberate choice rather than a silent gap. Always add at least one Alert Config (typically at minimum a Status-based one) unless you genuinely want a check that only records history without alerting.

Creating checks and Alert Configs together

The admin UI's check-creation form creates the check and any Alert Configs you've configured transactionally, in one step — if any Alert Config fails validation (e.g. an AlertFor not allowed for the check's type), the whole check creation is rolled back rather than leaving a check with no working alerts.

Occurrence tracking

Rather than creating a brand-new Alert record every time the same failing condition is re-observed, Piro increments OccurrenceCount on the existing open Alert. This keeps a single alert as the source of truth for "how many times has this been observed" instead of flooding the Alerts panel with duplicates for a condition that never actually resolved. MessageFingerprint is what identifies "the same alert" across occurrences.

Lifecycle

Fired → (re-fires increment OccurrenceCount) → Acknowledged (optional) → Resolved
  1. Fired — the alert config's failure threshold is met (or, for an External alert, the source system reports an open occurrence); an Alert record is created (or an existing unresolved one has its occurrence count bumped).
  2. Escalation (optional) — if an Escalation Policy is assigned (see Escalation snapshot), the alert begins progressing through its steps, paging on-call users until acknowledged.
  3. Acknowledged (optional) — a team member acknowledges the alert, which pauses escalation progression (subject to re-escalation thresholds). Acknowledging is blocked once an alert is Resolved — both the button is hidden in the admin UI and the backend rejects the action, since acknowledging something that's already over doesn't make sense.
  4. Resolved — the alert config's recovery threshold is met, or (for an External alert) the source system reports the occurrence closed; ResolvedAt is set and escalation stops.

The alert detail page in the admin UI shows an explicit Active/Resolved status badge, separate from the impact/severity pill. The severity pill is hidden once an alert is Resolved, to avoid a confusing "still Down" indicator sitting next to a "Resolved" badge.

Manual linking to an Incident

Alerts are never auto-attached to an Incident. From the Alerts panel, a user can:

  • Create a new incident from an alert — omit incidentId in the request and Piro creates one, pre-populated with the affected service.
  • Link to an existing incident — pass an incidentId to attach the alert to an incident already in progress (useful when multiple alerts represent the same outage).

For an External alert, there's no Service to pre-populate automatically, so the request also accepts a serviceIds list letting the admin explicitly pick which Service(s) — if any — the resulting Incident should be attached to. Picking zero services is valid: the Incident is simply created with no affected services, which the admin panel and public status page already handle gracefully ("no services listed").

Both paths go through the same endpoint:

POST /api/v1/alerts/{id}/incident

Source: AlertsOverviewController.cs

⚠️ No auto-creation. Earlier design notes/roadmaps for Piro mentioned automatically creating incidents from alerts — this was deliberately dropped. An alert firing does not by itself create or update any incident; someone must act on it from the Alerts panel.

Dashboard metrics

Alerts feed MTTA (mean time to acknowledge) and MTTR (mean time to resolve) metrics shown on the dashboard, computed from FiredAt, AcknowledgedAt, and ResolvedAt.

Related

  • Incidents — the user-facing disruption an alert can be manually linked to
  • Escalation Policies — how an unacknowledged alert escalates to on-call
  • Triggers — notification channels an alert dispatches to when it fires
  • Checks — the underlying probe whose result produces internal alerts
  • Integrations — inbound integrations (e.g. GCP Cloud Monitoring) that produce External alerts

Clone this wiki locally