docs / Feeding data in

Registering services

The one HTTP call that puts a service in the catalogue, every field it accepts, and where to keep it so it stays in Git.

A service exists in ITOps because something registered it. The registration is one JSON document sent to POST /api/v1/operator/register, and it carries everything about the service that cannot be observed from the outside: who owns it, which SLA group it belongs to, whether a backup is expected, what it depends on. Agents and push clients then report state; they never invent catalogue entries with metadata of their own.

Registration is idempotent. The same body sent twice updates the same service, identified by its path. That is what makes it safe to run from a Helm post-install hook, a CI step or a cron job, which is where it belongs: next to the thing it describes, versioned with it.

The path

Every service is identified by five levels:

organization / platform / environment / cluster / service
meridian     / commerce / prod        / eu-central-1 / payment-gateway

The first four levels are the node. An agent identifies itself with a node id, and asks for the services registered under it. The tree in the UI is this path; missing levels are created the first time a path is seen, so there is nothing to set up in advance. Choose stable, lowercase names: the path is the service's identity for as long as it reports.

The minimum

curl -X POST $API/api/v1/operator/register \
  -H "X-API-Key: $KEY" -H "Content-Type: application/json" -d '{
    "nodeId": "meridian/commerce/prod/eu-central-1",
    "name": "payment-gateway"
  }'

nodeId and name are the only required fields. name must be a valid Kubernetes resource name (lowercase RFC 1123 DNS subdomain), because for a Kubernetes service it is the workload name the agent will look for. The response:

{ "success": true, "serviceId": "3f9c…", "created": true }

created is false on an update.

The full body

Every field is optional beyond the two above. Each one unlocks something in the UI or in a plugin; the right-hand column says what.

Identity and presentation

Field Type Unlocks
nodeId string, required the four-level node the service hangs under
name string, required identity; for Kubernetes, the workload name
namespace string where the agent looks for the workload; defaults to the agent's namespace
displayName string what the UI shows instead of the name
description string shown on the service page and in search
serviceType string api, backend, frontend, worker, database, cache, queue, storage, network, application, external; drives the icon and filters
criticality string low, medium (default), high, critical; sorting, default SLA tier, ticket impact
tags string[] free-form filters in the catalogue
configSource string a URL to the Git location of this definition, linked from the service page
source string manual, operator, template; informational

SLA and incidents

Field Type Unlocks
slaGroup string membership in an SLA group. Groups are created on the fly; targets are set through the status sync (below) or the UI
slaDefinitionId string pin a specific SLA definition instead of the tier derived from criticality
autoIncident boolean open an SLA incident and a ticket when the service goes DOWN. Omitting the field leaves the stored value alone, so a re-registration cannot switch it off by accident
incidentCatalogItem string which catalogue item templates the auto-ticket; its workflow decides the queue. The default catalogue has outage_report
"ownership": {
  "team": "Payments", "owner": "Marcus Weber",
  "contact": "payments@example.com", "slack": "#payments",
  "onCall": "payments-oncall", "escalation": "Eszter Nagy"
},
"links": {
  "documentation": "https://…", "runbook": "https://…", "dashboard": "https://…",
  "repository": "https://…", "logs": "https://…", "alerts": "https://…", "api": "https://…",
  "custom": { "Vendor portal": "https://…" }
}

Ownership is shown on the service page and copied into auto-tickets. Every link becomes a button.

Dependencies

"dependencies": {
  "requires":   [ { "name": "postgres-primary", "nodePath": "meridian/commerce/prod/eu-central-1", "critical": true, "type": "database" } ],
  "requiredBy": [ { "name": "storefront", "nodePath": "meridian/commerce/prod/eu-central-1" } ]
},
"relations": { "parent": "commerce-platform", "children": [] }

One edge, declared on either side, is rendered in both directions: the service page shows what it needs and what needs it, and the CMDB page shows the whole graph. nodePath defaults to the service's own node, so dependencies inside one cluster need only a name. relations is a free-form logical hierarchy on top of the physical tree.

Operations and backups

"operations": {
  "tier": "tier1",
  "maintenanceWindow": "0 2 * * 0",
  "changeWindow": "0 9-17 * * 1-5",
  "desiredReplicas": 3, "minHealthyReplicas": 2,
  "backup": { "expected": true, "maxAgeDays": 1, "schedule": "0 1 * * *", "retention": "30d", "location": "s3://backups/orders" }
}

backup.expected with maxAgeDays is what turns silence into an alert: a service that expects a backup and has not reported a successful one within maxAgeDays is flagged overdue, and so is one that has never reported at all. Backup jobs report through POST /api/v1/backup/report.

Monitoring and health checks

"monitoring": { "enabled": true, "metricsPath": "/metrics", "metricsPort": 9090, "prometheusJob": "payment-gateway", "healthEndpoint": "/healthz" },
"healthCheck": { "enabled": true, "path": "/healthz", "port": 8080, "interval": "30s", "timeout": "5s" }

Stored and shown; the platform does not scrape Prometheus itself. HTTP probing is done by the agent's ITOPS_TARGETS (see the agent page), which registers its own targets.

Metadata

"metadata": {
  "version": "2.4.1", "port": 8080, "protocol": "http",
  "custom": { "costCenter": "CC-1042", "pciScope": true }
},
"custom": { "anything": "else" }

SLA groups and their targets

A service names its group with slaGroup; the group's display name, tier and targets are declared through the status sync endpoint, which an agent, a seed job or a curl can call:

curl -X POST $API/api/v1/operator/status \
  -H "X-API-Key: $KEY" -H "Content-Type: application/json" -d '{
    "nodeId": "meridian/commerce/prod/eu-central-1",
    "operatorVersion": "seed", "services": [],
    "slaGroups": [
      { "name": "checkout-flow",   "displayName": "Online checkout",       "tier": "critical", "targets": { "uptime": 99.95 } },
      { "name": "orders-database", "displayName": "Order database cluster", "tier": "critical", "targets": { "uptime": 99.99 } },
      { "name": "back-office",     "displayName": "Warehouse & back office", "tier": "high",   "targets": { "uptime": 99.5 } }
    ]
  }'

targets accepts uptime, responseTime and resolutionTime. A group without explicit targets inherits its tier's defaults, listed in SLA measurement.

Where to keep it

Helm post-install hook. A Job with helm.sh/hook: post-install,post-upgrade that runs the curl above, with the API key from a Secret. The registration then lives in the same chart as the service, is applied by the same pipeline, and updates when the chart does.

CI job. The deployment step that applies the manifests runs the same curl afterwards. Store the body as itops.json next to the manifests.

Seed job. For an estate you want to describe in one place, a script that loops over a list; the demo does exactly this with a ~200-line shell script and jq.

Whatever the vehicle, configSource should point at the file, so the service page links back to the definition that created it.

Services that register themselves

The three push endpoints, /api/v1/health/report, /api/v1/storage/report and /api/v1/backup/report, create the service on first sight when it does not exist, using the displayName, criticality, slaGroup, serviceType and tags in the push. That is enough for a host that only ever pushes. A full registration afterwards, with the same path, adds the rest without disturbing the history.

Reading the catalogue back

GET /api/v1/operator/services?nodeId=meridian/commerce/prod/eu-central-1 returns what the agent sees: id, the five-level externalId, name, namespace, displayName, criticality and sla. The GraphQL services and service queries return everything above, including ownership, links and both directions of the dependency graph.

Documentation for ITOps 4.2 · charts itops 2.0.0, sla-portal 1.4.0 · rendered 2026-09-11