-
Notifications
You must be signed in to change notification settings - Fork 2
Concepts Alerts
An Alert is the runtime record of something that fired — either a threshold rule (Alert Config) firing on a specific Check, or a notification received from a third-party system via an inbound Integration (currently: GCP Cloud Monitoring). It is a distinct entity from an Incident — an Alert tracks "this condition has been active since X"; an Incident is a manually declared, user-facing disruption. Alerts never create or attach to Incidents automatically — that link is always a deliberate action taken from the Alerts panel.
A check can have multiple Alert Configs. Since RFC 0002, Alert Configs are N:1 with Checks — one check can have several severity rules attached to it (for example, a "Warning" rule and a separate "Critical" rule for the same metric). See Alert Configs below.
| Field | Type | Description |
|---|---|---|
| AlertConfigId | int? | The threshold rule that fired. Null for an External alert — there is no internal Alert Config behind a third-party occurrence |
| CheckId | int? | The check whose result triggered this alert. Null for an External alert |
| ServiceId | int? | The service the check belongs to. Null for an External alert |
| IncidentId | int? | Set once this alert is linked to an incident (null until that manual step happens) |
| ImpactAtFireTime | ServiceStatus | The impact snapshot at the moment the alert fired — derived from the Alert Config's severity (internal) or mapped from the source's own severity (external) |
| Message | string? | Optional detail captured at fire time |
| MessageFingerprint | string | Dedup key for internal alerts — see Occurrence tracking |
| FiredAt | datetime | When the alert first fired |
| ResolvedAt | datetime? | When the alert resolved (recovery threshold met, or a closed event for an External alert). Null while still active |
| OccurrenceCount | int | How many times this same alert condition has re-fired without resolving — see below |
| EscalationCurrentStep / EscalationStepStartedAt | int?, datetime? | Tracks progress through the assigned Escalation Policy, if any |
| EscalationPolicyId | int? | The escalation policy snapshotted for this alert at creation time — see Escalation snapshot |
| AcknowledgedAt / AcknowledgedBy | long?, string? | Set when a team member acknowledges the alert. Blocked once the alert is Resolved — see Lifecycle |
| LastUserActivityAt | datetime? | Last time a human interacted with this alert — feeds escalation re-triggers |
| Source | enum | Where the alert came from — Internal (Piro's own check evaluation) or GcpCloudMonitoring. See External alerts
|
| ExternalId | string? | The source system's own identifier for this occurrence (e.g. GCP's incident_id). Null for Internal alerts, which dedup via MessageFingerprint instead |
| SourceUrl | string? | Optional deep link back into the source system's own console for this occurrence (e.g. GCP Cloud Monitoring's incident URL). Shown as a "View in GCP Cloud Monitoring"-style link on the alert detail page when present |
| SourceRequestLogId | int? | Points at the exact inbound webhook request that produced this alert, if any — see WebhookRequestLog |
Source: Alert.cs
Most alerts are anchored to a Check and a Service, as described above. An alert can also arrive with no Check and no Service at all — the admin UI calls this an External alert (an alert "managed externally of Piro"). This happens today only via the GCP Cloud Monitoring webhook integration: Piro receives a notification about something GCP's own alerting policy decided to fire, with nothing in Piro to correlate it against.
For an External alert:
-
AlertConfigId,CheckId, andServiceIdare all null. - Fields that only make sense for a threshold rule —
AlertFor,AlertValue, and severity thresholds — are not applicable. The admin UI shows "Managed externally of Piro" in the Trigger Criteria section instead of those fields. - Dedup uses
(Source, ExternalId)instead ofMessageFingerprint— repeated deliveries for the same source-side incident update the same Alert rather than creating duplicates. - It can still escalate to on-call — see Escalation snapshot below.
- It can still be promoted to an Incident, and the admin explicitly chooses which Service(s), if any, the resulting Incident should be attached to — see Manual linking to an Incident.
🔍 Correlating an incoming GCP alert to a specific Piro Check/Service is not implemented yet — every alert from the GCP Cloud Monitoring integration is currently External. This is a deliberate scope decision for the first version of that integration, not a limitation of the Alert model itself.
EscalationPolicyId on an Alert is set once, at creation time, from whichever policy applies:
- For a normal (Check-anchored) alert: the Service's
escalationPolicyId. - For an External alert: the source Integration's own
EscalationPolicyId, if the admin configured one on that integration.
This is a frozen snapshot, not a live reference — editing the policy afterward does not retroactively change the behavior of an alert that's already partway through escalation.
An Alert Config is the threshold rule attached to a check. As of RFC 0002, checks no longer judge their own severity — a check's executor only reports whether it measured successfully (UP) or failed to measure at all (DOWN, e.g. connection refused, DNS didn't resolve, TLS handshake failed), plus a raw metric value (latency in ms, days until certificate expiry, count of failed name servers, etc.). Deciding what that raw value means — "this counts as degraded," "this counts as critical," and at what threshold — is entirely the Alert Config's job.
A single check can have any number of Alert Configs, each independently evaluating the check's raw metric. This lets you layer severities on the same signal instead of picking one threshold — for example, on an SSL check:
| Alert Config | AlertFor | AlertValue | Severity |
|---|---|---|---|
| "Warning — renew soon" | CertExpiry |
30 (days) |
Warning |
| "Critical — expiring imminently" | CertExpiry |
7 (days) |
Critical |
Both configs watch the same MetricValue (days remaining) reported by the check; each fires independently once the certificate crosses its own threshold. Previously (pre-RFC 0002) a check could only have one Alert Config — this is the key behavior change to be aware of when migrating existing checks.
| Field | Type | Description |
|---|---|---|
| CheckId | int | The check this rule is attached to |
| AlertFor | enum | The metric this rule evaluates — see below |
| AlertValue | string | The threshold: a ServiceStatus name for Status alerts, milliseconds for Latency, days for CertExpiry, or a failed-name-server count for FailedNameServers
|
| Severity | enum |
Warning or Critical
|
| FailureThreshold | int | Consecutive failing measurements before the alert fires (default 1) |
| SuccessThreshold | int | Consecutive passing measurements before the alert auto-resolves (default 1) |
| IsActive | bool | Disables the rule without deleting it |
| Description | string? | Optional free-text note |
Source: AlertConfig.cs
| AlertFor | Compares against | Meaning |
|---|---|---|
Status |
the check's UP/DOWN result |
Fires based on the executor's own success/failure determination |
Latency |
CheckDataPoint.LatencyMs |
Fires when response/resolution time crosses AlertValue (ms) |
CertExpiry |
CheckDataPoint.MetricValue |
Fires when days-until-expiry drops to or below AlertValue. SSL checks only |
FailedNameServers |
CheckDataPoint.MetricValue |
Fires when the count of name servers that failed to resolve reaches AlertValue. DNS checks only |
🔍 The old
UptimeAlertFor was removed — it was never actually implemented and has no replacement; useStatusfor availability-based alerting.
Source: AlertFor.cs, AlertSeverity.cs
Not every AlertFor makes sense for every Check type — a GCP Cloud Run Job check has no latency signal, and only SSL checks report certificate expiry. The admin UI only offers the values valid for the check's type, enforced server-side:
| CheckType | Allowed AlertFor values |
|---|---|
HTTP |
Status, Latency
|
DNS |
Status, Latency, FailedNameServers
|
TCP |
Status, Latency
|
Ping |
Status, Latency
|
SSL |
Status, CertExpiry
|
Heartbeat |
Status |
GRPC |
Status, Latency
|
GCP_CloudRunJob |
Status |
Source: CheckTypeExtensions.cs
Creating a check does not create a default Alert Config. If you save a check without adding at least one Alert Config, it will run on schedule and record data, but nothing will ever notify you — there is no built-in fallback severity rule.
⚠️ The admin panel warns you about this. When creating a check with zero Alert Configs configured, the check-creation form shows a confirmation dialog before letting you proceed, so this is a deliberate choice rather than a silent gap. Always add at least one Alert Config (typically at minimum aStatus-based one) unless you genuinely want a check that only records history without alerting.
The admin UI's check-creation form creates the check and any Alert Configs you've configured transactionally, in one step — if any Alert Config fails validation (e.g. an AlertFor not allowed for the check's type), the whole check creation is rolled back rather than leaving a check with no working alerts.
Rather than creating a brand-new Alert record every time the same failing condition is re-observed, Piro increments OccurrenceCount on the existing open Alert. This keeps a single alert as the source of truth for "how many times has this been observed" instead of flooding the Alerts panel with duplicates for a condition that never actually resolved. MessageFingerprint is what identifies "the same alert" across occurrences.
Fired → (re-fires increment OccurrenceCount) → Acknowledged (optional) → Resolved
- Fired — the alert config's failure threshold is met (or, for an External alert, the source system reports an open occurrence); an Alert record is created (or an existing unresolved one has its occurrence count bumped).
- Escalation (optional) — if an Escalation Policy is assigned (see Escalation snapshot), the alert begins progressing through its steps, paging on-call users until acknowledged.
- Acknowledged (optional) — a team member acknowledges the alert, which pauses escalation progression (subject to re-escalation thresholds). Acknowledging is blocked once an alert is Resolved — both the button is hidden in the admin UI and the backend rejects the action, since acknowledging something that's already over doesn't make sense.
-
Resolved — the alert config's recovery threshold is met, or (for an External alert) the source system reports the occurrence closed;
ResolvedAtis set and escalation stops.
The alert detail page in the admin UI shows an explicit Active/Resolved status badge, separate from the impact/severity pill. The severity pill is hidden once an alert is Resolved, to avoid a confusing "still Down" indicator sitting next to a "Resolved" badge.
Alerts are never auto-attached to an Incident. From the Alerts panel, a user can:
-
Create a new incident from an alert — omit
incidentIdin the request and Piro creates one, pre-populated with the affected service. -
Link to an existing incident — pass an
incidentIdto attach the alert to an incident already in progress (useful when multiple alerts represent the same outage).
For an External alert, there's no Service to pre-populate automatically, so the request also accepts a serviceIds list letting the admin explicitly pick which Service(s) — if any — the resulting Incident should be attached to. Picking zero services is valid: the Incident is simply created with no affected services, which the admin panel and public status page already handle gracefully ("no services listed").
Both paths go through the same endpoint:
POST /api/v1/alerts/{id}/incident
Source: AlertsOverviewController.cs
⚠️ No auto-creation. Earlier design notes/roadmaps for Piro mentioned automatically creating incidents from alerts — this was deliberately dropped. An alert firing does not by itself create or update any incident; someone must act on it from the Alerts panel.
Alerts feed MTTA (mean time to acknowledge) and MTTR (mean time to resolve) metrics shown on the dashboard, computed from FiredAt, AcknowledgedAt, and ResolvedAt.
- Incidents — the user-facing disruption an alert can be manually linked to
- Escalation Policies — how an unacknowledged alert escalates to on-call
- Triggers — notification channels an alert dispatches to when it fires
- Checks — the underlying probe whose result produces internal alerts
- Integrations — inbound integrations (e.g. GCP Cloud Monitoring) that produce External alerts