Overview
What ITOps is, how the pieces fit, and where to go next.
ITOps is a self-hosted IT operations platform. It keeps a catalogue of the services you run, measures their uptime against SLA targets, opens and tracks tickets when something breaks, and publishes a status page your customers can read. It runs in your own Kubernetes cluster and it watches Kubernetes workloads, virtual machines, bare metal, appliances and third-party services with the same model.
Monitoring tools tell you what is happening. ITOps is the layer above them: what it means, who owns it, and what should happen next.
The pieces
| Component | What it is | Where it runs |
|---|---|---|
| core | Go backend: GraphQL and REST API, the SLA and ticketing plugins, PostgreSQL storage | your cluster, itops chart |
| ui | Vue single-page app, 10 languages | your cluster, itops chart |
| itops-agentv2 | ~1,100 lines of Go, stdlib only, read-only. Reports workload state and probes HTTP endpoints | any namespace you want watched, or anywhere with network access for HTTP probes |
| itops-agent-sh | the same heartbeat in ~200 lines of POSIX shell, for hosts that only have curl |
VMs, bare metal, appliances |
| sla-portal | standalone public status page with its own SQLite store | your cluster, sla-portal chart, or anywhere reachable from core |
The core is free. The SLA and ticketing plugins are activated by a licence key; without one the platform runs the catalogue, the agents, the push endpoints, CMDB, webhooks and access control, and nothing else.
Architecture in sixty seconds
you / CI / Helm hook ──── POST /api/v1/operator/register ───┐
▼
itops-agentv2 (k8s) ──── GET /api/v1/operator/services ── core ──▶ PostgreSQL
│ POST /api/v1/operator/status │
│ POST /api/v1/operator/heartbeat │ daily report
└── HTTP probes ── POST /api/v1/health/report ▼
sla-portal ──▶ public status page
VM / cron / backup job ─ POST /api/v1/health|storage|backup/report
│
browser ──────────────── /graphql, /graphql/ws ─────────────┘
Two kinds of information reach the core, and it is worth keeping them apart:
- What should exist. Services are registered through
POST /api/v1/operator/register, one JSON body per service. That body carries everything the platform will ever know about the service that cannot be observed: its owner, its SLA group, whether a backup is expected, what it depends on. Registration is idempotent, so it belongs in a Helm hook or a CI job next to the service it describes. - What does exist. Agents and push clients report state. The Kubernetes agent asks the core which services it should look for and reports replica counts. Everything else, from a bare-metal database to a SaaS endpoint, reports through the three push endpoints.
Every service is identified by one five-level path, organization/platform/environment/cluster/service, for example meridian/commerce/prod/eu-central-1/payment-gateway. The first four levels are the node, which is what an agent identifies itself as. The tree in the UI is this path, and it is created on demand the first time a path is seen.
What the SLA numbers are made of
Every service's status is sampled into a snapshot every five minutes. Uptime for a service, and for a group of services, is computed from those snapshots on an exact-duration timeline, weighted by minutes, with maintenance windows removed from both the numerator and the denominator. Time the platform could not measure is reported as UNKNOWN rather than being counted as up. Each SLA group has a target, an error budget expressed in minutes, and a projection that flags the month as AT_RISK before the budget runs out. The details are in SLA measurement.
Where to go next
- Quickstart: a running platform and a measured service in ten minutes.
- Install: the production values, external PostgreSQL, ingress, TLS and network policy.
- Registering services: the full registration schema, one field at a time.
- Kubernetes agent and VMs and bare metal: getting state into the platform.
- API reference: every endpoint, the GraphQL map and the WebSocket events.
- Prompt files for AI assistants: if you would rather have Claude or another assistant write the YAML.
Getting the source of the docs
These pages are rendered from Markdown in the apps/landing/docs-src directory of the product repository. If a page is wrong, the fix is a pull request.