Push API
The three endpoints anything can report to - health, storage and backups - with their payloads, what the core does with them, and what happens when the reports stop.
Three endpoints accept reports from anything that can make an HTTP request: a cron job, a backup tool's post-hook, a CI pipeline, a monitoring system's webhook. They share the same shape: authenticate with the operator API key, name the service by its path, say what happened. A service that does not exist yet is created on first sight, so nothing has to be registered in advance.
All three:
POST,Content-Type: application/json, body up to 64 KiBX-API-Key: <ITOPS_SECURITY_OPERATOR_API_KEY>(orAuthorization: Bearer <key>)- identify the service with
path(five levels) or withnodeId+service - answer with
{"success": true, …, "warnings": [...]}; a warning means the report was accepted but something in it was adjusted
The path shortcut
{ "path": "acme/shop/prod/dc-budapest/galera-node-1" }
is the same as
{ "nodeId": "acme/shop/prod/dc-budapest", "service": "galera-node-1" }
A nodeId with fewer than four segments is padded with the literal unknown and a warning is returned, rather than the report being rejected: the data lands, visibly, under an unknown/… branch you can then fix. Values that are too long or contain control characters are trimmed, with a warning.
Health
POST /api/v1/health/report
{
"path": "acme/shop/prod/dc-budapest/galera-node-1",
"status": "OPERATIONAL",
"message": "wsrep_cluster_size=3, wsrep_ready=ON",
"displayName": "Galera node 1",
"criticality": "critical",
"slaGroup": "orders-database",
"serviceType": "database",
"tags": ["galera", "baremetal"],
"metadata": { "rack": "B04" }
}
| Field | Notes |
|---|---|
status |
OPERATIONAL, DEGRADED, DOWN, MAINTENANCE, UNKNOWN. Anything else, or missing, becomes UNKNOWN with a warning. |
message |
Free text, shown next to the status and kept in the status history. Put the evidence here. |
displayName, criticality, slaGroup, serviceType, tags, metadata |
Used when the service is created; ignored on later reports. Change them with a registration. |
Response: {"success":true,"serviceId":"…","created":false,"status":"OPERATIONAL","warnings":[]}.
On first creation the service gets the SLA tier its criticality implies. A status change to DOWN opens an SLA incident and, if the service has autoIncident, a ticket; a change back closes them. See SLA measurement and Ticketing.
Storage
POST /api/v1/storage/report
{
"path": "acme/shop/prod/dc-budapest/galera-node-1",
"allocatedBytes": 536870912000,
"usedBytes": 402653184000,
"storageType": "disk",
"mountPath": "/var/lib/mysql",
"message": "nightly df",
"reportedBy": "cron@galera-node-1"
}
Send allocatedBytes and usedBytes, or freePercent directly. The core computes the free percentage and a status:
| Free | Status |
|---|---|
| below 10 % | critical |
| below 30 % | warning |
| otherwise | healthy |
storageType is disk, pvc, s3 or rds, for the icon. Out-of-range values are clamped with a warning. The Storage tab shows the latest report per service and the trend.
Backups
POST /api/v1/backup/report
{
"path": "acme/shop/prod/dc-budapest/galera-node-1",
"status": "success",
"sizeBytes": 734003200,
"message": "xtrabackup full, 11m42s",
"reportedBy": "backup-cron"
}
status is success, failed or partial. The target is one of, in this order of precedence: namespace (every service in it), slaGroup (every member), or a single service by path or service. The first report sets backup.expected on the service; from then on the service is judged.
Overdue is what makes this useful. A service whose registration says operations.backup.expected: true with maxAgeDays: N is flagged overdue when its last successful backup is older than N days, or when it has never reported a backup at all. The flag is on the service page and in the Backup tab; a failed report also fires the backup.failed webhook event. Silence is treated as failure, which is the only honest reading of a backup job that did not run.
Make the report the last line of whatever runs your backups:
if velero backup create nightly-$(date +%F) --wait; then s=success; else s=failed; fi
curl -sS -X POST "$ITOPS_URL/api/v1/backup/report" -H "X-API-Key: $KEY" \
-H "Content-Type: application/json" \
-d "{\"namespace\":\"shop\",\"status\":\"$s\",\"reportedBy\":\"velero\"}"
When the reports stop
A background check runs every 60 seconds over services whose source is external, that is, services created by a push rather than by the Kubernetes agent:
| Silent for | Becomes | Message |
|---|---|---|
| more than 2 minutes | UNKNOWN |
No heartbeat received |
| more than 5 minutes | OUTAGE |
Service not reporting (stale) |
The outage check runs first, so a service that has been silent for an hour lands in OUTAGE and does not stick at UNKNOWN. MAINTENANCE is never overwritten by either check. Choose the push interval accordingly: once a minute is the usual cron cadence, and it leaves one missed run before the status changes.
Authentication failures
A rejected key is answered with 401 and written to the SLA event log as auth_failed with the method and path, so a misconfigured host is visible from the dashboard rather than only from the host's own logs.
The other operator endpoints
The agent's own endpoints, GET /api/v1/operator/services, POST /api/v1/operator/status, POST /api/v1/operator/heartbeat and POST /api/v1/operator/register, use the same key and are documented on the API reference. The status sync is also how SLA groups get their targets.