docs / Feeding data in

Push API

The three endpoints anything can report to - health, storage and backups - with their payloads, what the core does with them, and what happens when the reports stop.

Three endpoints accept reports from anything that can make an HTTP request: a cron job, a backup tool's post-hook, a CI pipeline, a monitoring system's webhook. They share the same shape: authenticate with the operator API key, name the service by its path, say what happened. A service that does not exist yet is created on first sight, so nothing has to be registered in advance.

All three:

  • POST, Content-Type: application/json, body up to 64 KiB
  • X-API-Key: <ITOPS_SECURITY_OPERATOR_API_KEY> (or Authorization: Bearer <key>)
  • identify the service with path (five levels) or with nodeId + service
  • answer with {"success": true, …, "warnings": [...]}; a warning means the report was accepted but something in it was adjusted

The path shortcut

{ "path": "acme/shop/prod/dc-budapest/galera-node-1" }

is the same as

{ "nodeId": "acme/shop/prod/dc-budapest", "service": "galera-node-1" }

A nodeId with fewer than four segments is padded with the literal unknown and a warning is returned, rather than the report being rejected: the data lands, visibly, under an unknown/… branch you can then fix. Values that are too long or contain control characters are trimmed, with a warning.

Health

POST /api/v1/health/report

{
  "path": "acme/shop/prod/dc-budapest/galera-node-1",
  "status": "OPERATIONAL",
  "message": "wsrep_cluster_size=3, wsrep_ready=ON",
  "displayName": "Galera node 1",
  "criticality": "critical",
  "slaGroup": "orders-database",
  "serviceType": "database",
  "tags": ["galera", "baremetal"],
  "metadata": { "rack": "B04" }
}
Field Notes
status OPERATIONAL, DEGRADED, DOWN, MAINTENANCE, UNKNOWN. Anything else, or missing, becomes UNKNOWN with a warning.
message Free text, shown next to the status and kept in the status history. Put the evidence here.
displayName, criticality, slaGroup, serviceType, tags, metadata Used when the service is created; ignored on later reports. Change them with a registration.

Response: {"success":true,"serviceId":"…","created":false,"status":"OPERATIONAL","warnings":[]}.

On first creation the service gets the SLA tier its criticality implies. A status change to DOWN opens an SLA incident and, if the service has autoIncident, a ticket; a change back closes them. See SLA measurement and Ticketing.

Storage

POST /api/v1/storage/report

{
  "path": "acme/shop/prod/dc-budapest/galera-node-1",
  "allocatedBytes": 536870912000,
  "usedBytes": 402653184000,
  "storageType": "disk",
  "mountPath": "/var/lib/mysql",
  "message": "nightly df",
  "reportedBy": "cron@galera-node-1"
}

Send allocatedBytes and usedBytes, or freePercent directly. The core computes the free percentage and a status:

Free Status
below 10 % critical
below 30 % warning
otherwise healthy

storageType is disk, pvc, s3 or rds, for the icon. Out-of-range values are clamped with a warning. The Storage tab shows the latest report per service and the trend.

Backups

POST /api/v1/backup/report

{
  "path": "acme/shop/prod/dc-budapest/galera-node-1",
  "status": "success",
  "sizeBytes": 734003200,
  "message": "xtrabackup full, 11m42s",
  "reportedBy": "backup-cron"
}

status is success, failed or partial. The target is one of, in this order of precedence: namespace (every service in it), slaGroup (every member), or a single service by path or service. The first report sets backup.expected on the service; from then on the service is judged.

Overdue is what makes this useful. A service whose registration says operations.backup.expected: true with maxAgeDays: N is flagged overdue when its last successful backup is older than N days, or when it has never reported a backup at all. The flag is on the service page and in the Backup tab; a failed report also fires the backup.failed webhook event. Silence is treated as failure, which is the only honest reading of a backup job that did not run.

Make the report the last line of whatever runs your backups:

if velero backup create nightly-$(date +%F) --wait; then s=success; else s=failed; fi
curl -sS -X POST "$ITOPS_URL/api/v1/backup/report" -H "X-API-Key: $KEY" \
  -H "Content-Type: application/json" \
  -d "{\"namespace\":\"shop\",\"status\":\"$s\",\"reportedBy\":\"velero\"}"

When the reports stop

A background check runs every 60 seconds over services whose source is external, that is, services created by a push rather than by the Kubernetes agent:

Silent for Becomes Message
more than 2 minutes UNKNOWN No heartbeat received
more than 5 minutes OUTAGE Service not reporting (stale)

The outage check runs first, so a service that has been silent for an hour lands in OUTAGE and does not stick at UNKNOWN. MAINTENANCE is never overwritten by either check. Choose the push interval accordingly: once a minute is the usual cron cadence, and it leaves one missed run before the status changes.

Authentication failures

A rejected key is answered with 401 and written to the SLA event log as auth_failed with the method and path, so a misconfigured host is visible from the dashboard rather than only from the host's own logs.

The other operator endpoints

The agent's own endpoints, GET /api/v1/operator/services, POST /api/v1/operator/status, POST /api/v1/operator/heartbeat and POST /api/v1/operator/register, use the same key and are documented on the API reference. The status sync is also how SLA groups get their targets.

Documentation for ITOps 4.2 · charts itops 2.0.0, sla-portal 1.4.0 · rendered 2026-09-11