SLA measurement
How uptime is measured, what an error budget is made of, why UNKNOWN stays UNKNOWN, and how maintenance windows and the daily report work.
The SLA plugin turns status history into numbers you can put in a contract: uptime per service and per group, an error budget in minutes, a projection for the month in flight, and a daily report. It is a licensed plugin; without a licence the snapshots still accumulate and the SLA screens are hidden, so nothing is lost by adding the key later.
Snapshots
Every five minutes the platform records the status of every service: OPERATIONAL, DEGRADED, DOWN, MAINTENANCE or UNKNOWN, with the timestamp of the status change that produced it. The aggregator runs every five minutes with a fifteen-minute delay, so that every agent and push client has had time to report before a window is closed.
Uptime is not "snapshots that were green divided by snapshots taken". It is computed on an exact-duration timeline: each status change is an interval with a start and an end, and the fraction is minutes up over minutes measured. A service that was down for 90 seconds contributes 90 seconds of downtime, not one five-minute bucket.
The fraction
For a service:
minutes OPERATIONAL (+ DEGRADED, weighted)
uptime = ───────────────────────────────────────────────────────────────
minutes measured − minutes in MAINTENANCE − minutes UNKNOWN
- Maintenance is removed from both sides. A declared window neither counts against you nor pads the denominator.
- Unknown is removed from both sides too, and reported separately. A day where the platform could not see the service for six hours shows
uptime 99.8 %, unknown 25 %, notuptime 100 %. A day with too little measured time to be meaningful is reported asUNKNOWNrather thanMET. - Degraded costs half a minute of downtime per minute by default (weight 0.5). A service answering slowly is not the same as one not answering at all, and a contract that cannot express the difference pushes both sides to argue about labels.
For a group, services are weighted by minutes measured. A large service that was watched all month moves the headline number; a service added on the 28th does not.
Tiers
Four SLA definitions are seeded and assigned by criticality when a service is created. A group's targets override the uptime figure; a service's slaDefinitionId pins a definition explicitly.
| Tier | Assigned to criticality | Uptime | Response | Resolution | Hours |
|---|---|---|---|---|---|
sla-tier-critical |
critical |
99.99 % | 15 min | 4 h | 24×7 |
sla-tier-high |
high |
99.9 % | 60 min | 8 h | 24×7 |
sla-tier-medium |
medium |
99.5 % | 4 h | 72 h | business hours |
sla-tier-low |
low |
99.0 % | 24 h | 120 h | business hours |
Business hours default to 08:00–18:00 Monday to Friday in ITOPS_COMPANY_TIMEZONE and are DST-safe. The response and resolution targets are the same table the ticketing plugin uses for priority-driven deadlines, so the promise a queue is sorted by and the promise a report is judged against come from one place.
Groups
Services belong to SLA groups: checkout-flow, orders-database, back-office. A group is the unit customers and management care about, and the unit the status page shows. Membership is the slaGroup field on registration or on the first push; display name, tier and targets come from the status sync or the UI.
Every group shows, for the current month:
- uptime so far, against its target;
- error budget in minutes: allowed, used, remaining, and the burn rate;
- status:
MET,AT_RISKorBREACHED.AT_RISKis a projection: at the current burn rate, the budget will not last the month. It is the number to act on, because it arrives while there is still time to.
The error budget for a month is (1 − target) × minutes in the month. At 99.9 % on a 30-day month that is 43.2 minutes; at 99.99 % it is 4.3.
Incidents
An SLA incident opens when a service enters DOWN (or DEGRADED, per definition) and closes when it recovers. Incidents have a source (MONITORING for automatic ones, MANUAL for ones created in the UI), a severity, and a timeline. If the service has autoIncident, the ticketing plugin opens a ticket alongside; see Automatic incidents.
Every incident is on the service page, in the group's month view, in the daily report and, for the groups you publish, on the status page.
Maintenance windows
A maintenance window, called an exclusion window in the API, removes an interval from the measurement. Open one before a planned change and close it after:
curl -X POST $API/api/v1/sla/exclusion-window/start -H "X-API-Key: $KEY" \
-H "Content-Type: application/json" -d '{"serviceId": "<uuid>", "reason": "PostgreSQL major upgrade"}'
# → {"windowId": "…"}
curl -X POST $API/api/v1/sla/exclusion-window/stop -H "X-API-Key: $KEY" \
-H "Content-Type: application/json" -d '{"windowId": "…"}'
The same is available as the createExclusionWindow and stopExclusionWindow GraphQL mutations, and in the UI. The service's registration can also declare a recurring operations.maintenanceWindow in cron syntax, which is shown on the service page. Pushing MAINTENANCE as the status has the same effect on the fraction and is never overwritten by the stale check.
Windows are recorded, listed (slaExclusionWindows) and published in the daily report and on the status page, so a suspiciously clean month can be audited.
The daily report
Once a day the plugin generates a report per group: uptime, error budget, incidents with their durations, maintenance windows, and one line per service. It is stored as JSON, rendered to PDF, and pushed to the public status page if one is configured. Generate one by hand:
curl -X POST $API/api/v1/sla/report/generate -H "X-API-Key: $KEY"
The plugin keeps an event log (slaEventLog in GraphQL): sla_breached, data_gap, report_generated, report_failed, aggregation_completed and auth_failed, each with a timestamp and a payload. data_gap is the one to watch: it means the platform knows it was blind for a while.
Alerts
slaAlerts lists breaches and at-risk transitions; acknowledgeSLAAlert marks one as seen. The same transitions go out as the sla.breached, sla.at_risk and sla.met webhook events, and sla.daily_report fires when a report is generated. Burn-rate paging rules are not part of this release: the budget and the burn rate are computed and shown, and the alerting on them is yours to wire to a webhook.
GraphQL entry points
slaGroups, slaGroup, slaDefinitions, slaIncidents, slaPeriodResults, slaSnapshotTrend, slaTrendData, slaDashboardStats, slaAlerts, slaExclusionWindows, slaEventLog; mutations createSLADefinition, updateSLADefinition, createSLATarget, assignSLAToService, createSLAIncident, updateSLAIncident, createExclusionWindow, stopExclusionWindow, acknowledgeSLAAlert. The API reference has the shape of each.
Export and import
GET /api/v1/templates/export returns the SLA definitions as YAML; POST /api/v1/templates/import loads them. Use it to carry a tuned set of definitions from staging to production, or to keep them in Git.