docs / Start here

Overview

What ITOps is, how the pieces fit, and where to go next.

ITOps is a self-hosted IT operations platform. It keeps a catalogue of the services you run, measures their uptime against SLA targets, opens and tracks tickets when something breaks, and publishes a status page your customers can read. It runs in your own Kubernetes cluster and it watches Kubernetes workloads, virtual machines, bare metal, appliances and third-party services with the same model.

Monitoring tools tell you what is happening. ITOps is the layer above them: what it means, who owns it, and what should happen next.

The pieces

Component What it is Where it runs
core Go backend: GraphQL and REST API, the SLA and ticketing plugins, PostgreSQL storage your cluster, itops chart
ui Vue single-page app, 10 languages your cluster, itops chart
itops-agentv2 ~1,100 lines of Go, stdlib only, read-only. Reports workload state and probes HTTP endpoints any namespace you want watched, or anywhere with network access for HTTP probes
itops-agent-sh the same heartbeat in ~200 lines of POSIX shell, for hosts that only have curl VMs, bare metal, appliances
sla-portal standalone public status page with its own SQLite store your cluster, sla-portal chart, or anywhere reachable from core

The core is free. The SLA and ticketing plugins are activated by a licence key; without one the platform runs the catalogue, the agents, the push endpoints, CMDB, webhooks and access control, and nothing else.

Architecture in sixty seconds

  you / CI / Helm hook ──── POST /api/v1/operator/register ───┐
                                                              ▼
  itops-agentv2 (k8s) ──── GET  /api/v1/operator/services ── core ──▶ PostgreSQL
        │                  POST /api/v1/operator/status       │
        │                  POST /api/v1/operator/heartbeat     │ daily report
        └── HTTP probes ── POST /api/v1/health/report          ▼
                                                            sla-portal ──▶ public status page
  VM / cron / backup job ─ POST /api/v1/health|storage|backup/report
                                                              │
  browser ──────────────── /graphql, /graphql/ws ─────────────┘

Two kinds of information reach the core, and it is worth keeping them apart:

  1. What should exist. Services are registered through POST /api/v1/operator/register, one JSON body per service. That body carries everything the platform will ever know about the service that cannot be observed: its owner, its SLA group, whether a backup is expected, what it depends on. Registration is idempotent, so it belongs in a Helm hook or a CI job next to the service it describes.
  2. What does exist. Agents and push clients report state. The Kubernetes agent asks the core which services it should look for and reports replica counts. Everything else, from a bare-metal database to a SaaS endpoint, reports through the three push endpoints.

Every service is identified by one five-level path, organization/platform/environment/cluster/service, for example meridian/commerce/prod/eu-central-1/payment-gateway. The first four levels are the node, which is what an agent identifies itself as. The tree in the UI is this path, and it is created on demand the first time a path is seen.

What the SLA numbers are made of

Every service's status is sampled into a snapshot every five minutes. Uptime for a service, and for a group of services, is computed from those snapshots on an exact-duration timeline, weighted by minutes, with maintenance windows removed from both the numerator and the denominator. Time the platform could not measure is reported as UNKNOWN rather than being counted as up. Each SLA group has a target, an error budget expressed in minutes, and a projection that flags the month as AT_RISK before the budget runs out. The details are in SLA measurement.

Where to go next

Getting the source of the docs

These pages are rendered from Markdown in the apps/landing/docs-src directory of the product repository. If a page is wrong, the fix is a pull request.

Documentation for ITOps 4.2 · charts itops 2.0.0, sla-portal 1.4.0 · rendered 2026-09-11