// v4.2 — kubernetes-native it operations, self-hosted

Know your uptime.
Prove your SLA.

An IT operations platform that runs on your own cluster, for your Kubernetes workloads and for everything running next to them. One Helm install, and uptime numbers your board can read: measured, not estimated.

The demo runs Meridian Retail Group, a fictional online retailer: a Kubernetes commerce platform, a bare-metal database cluster and a few cloud services. Sign in with demo / MeridianDemo2026. The account reads everything and changes nothing.

itops — from zero to a measured service
$ helm repo add itops https://charts.mlops.hu
$ helm install itops itops/itops -n itops --create-namespace
NAME: itops · STATUS: deployed · REVISION: 1

$ curl -X POST $API/api/v1/operator/register \
    -H "X-API-Key: $KEY" -d '{
      "nodeId": "meridian/commerce/prod/eu-central-1",
      "name": "payment-gateway", "criticality": "critical",
      "slaGroup": "checkout-flow", "autoIncident": true }'
{"success":true,"created":true}

# 30 seconds later the agent reports in
✓ payment-gateway   OPERATIONAL   3/3 ready   99.99%   budget 92%
$ 
// the result, on one page

This is what your customers see

The public status page of the demo estate, as the SLA plugin publishes it: 90 days per group, the month's uptime against its target, the error budget left. When the platform was blind, the day is grey, not green.

All systems operational · Meridian Retail Groupupdated 12:35 UTC · last 90 days
Online checkoutcheckout-flow · target 99.95%
99.98%
budget 81%
Order databaseorders-database · target 99.99%
99.99%
budget 92%
Warehouse & back officeback-office · target 99.5%
99.61%
budget 12% · at risk
Shared platform servicesplatform-services · target 99.0%
99.87%
budget 88%
met degraded incident not measured open the live status page →
5 minSLA snapshot resolution
30 sagent report cycle
0third-party code in the agent
10languages in the UI
// features

Everything from the service catalogue to the SLA report your customer reads

One self-hosted platform. The core is free; the SLA and ticketing plugins unlock with a licence key.

sla/measurement

SLA measurement

Uptime is built from 5-minute snapshots on an exact-duration timeline and weighted by minutes, so a large service moves the headline number and a small one does not. Maintenance windows are excluded from both sides of the fraction. Time the platform could not measure is reported as UNKNOWN, never as a silent 100%.

sla/error-budget

Error budget, projected forward

Every SLA group shows its budget in minutes: allowed, used, remaining, and the burn rate. A month in flight is flagged AT_RISK before the budget is spent, not after the report is printed. Four seeded tiers from 99.0% to 99.99%, with your own targets per group.

ticketing/itil

ITIL ticketing

Incidents, service requests, problems and changes, kept apart so MTTR and request fulfilment stop blurring into one number. Priority is derived from impact × urgency, and the priority sets the deadlines: a P1 is answered in 15 minutes and resolved in 4 hours. Putting a ticket on hold stops the clock.

incidents/auto

Automatic incidents, automatic closure

When a service goes DOWN, the platform opens an SLA incident and a ticket in the right queue, with type, impact and urgency filled in. When the service comes back and nobody has touched the ticket, it resolves itself and says why.

agent/kubernetes

An agent you can read before you install it

itops-agentv2 is about 1,100 lines of Go with no third-party code, shipped in a scratch image. It reports replica counts and images every 30 seconds, needs get on workloads and nothing else, opens no port and never writes to your cluster. MIT licensed.

beyond/kubernetes

VMs, bare metal, appliances, SaaS

The same agent probes HTTP endpoints with no cluster at all. A 200-line POSIX shell agent covers hosts that only have curl. Anything that can send one HTTP request lands on the same dashboard with the same SLA maths, and a service that stops reporting turns UNKNOWN after 2 minutes and OUTAGE after 5.

status/public

Public status page

A standalone page for customers and management: uptime per SLA group, error budget left, incident history and maintenance windows for the whole month. No login, just a link. When the data is thin, it says so instead of showing green.

cmdb/dependencies

Catalogue, ownership, dependencies

Every service carries its team, owner, on-call, runbook and dashboard links, tags and tier. Dependencies are declared once and rendered in both directions, so "what breaks if this goes down" is one query, not a meeting.

backups/webhooks/rbac

Backups, storage, webhooks, access control

Any backup tool ends with one HTTP call, and an overdue backup raises an alert. Storage reports turn into healthy / warning / critical at 30% and 10% free. 25 event types go out as HMAC-signed webhooks with SSRF protection on every hop. Four roles from group membership, a read-only viewer role, field-level permissions, LDAP and Active Directory sign-in.

// how it works

Install, register, watch. Three steps, no admin clicking.

The catalogue belongs to the server. You tell it what should exist; the agents tell it what does. Both are plain HTTP calls you can put in a Helm hook, a CI job or a cron line.

01 / helm

Install into your cluster

Core, UI and a bundled PostgreSQL from one chart. Bring your own database, ingress and LDAP when you are ready for production.

$ helm repo add itops https://charts.mlops.hu
$ helm install itops itops/itops \
    -n itops --create-namespace
first login: admin / Password123!
change it. it is on the checklist.
02 / register

Say what should exist

Every service is identified by one path: org/platform/env/cluster/service. The name and the node are the only required fields. Everything else unlocks one more feature at a time.

$ curl -X POST $API/api/v1/operator/register \
   -H "X-API-Key: $KEY" -d '{
     "nodeId": "acme/shop/prod/eu-1",
     "name": "orders-db",
     "slaGroup": "orders-database",
     "operations": { "backup":
       { "expected": true, "maxAgeDays": 1 } }
   }'
03 / report

Let the estate report in

The Kubernetes agent looks each registered service up and reports its state. Hosts push their own health. Backup jobs end with one more request.

# any backup script, last line
$ curl -X POST $API/api/v1/backup/report \
   -H "X-API-Key: $KEY" -d '{
     "path": "acme/shop/prod/eu-1/orders-db",
     "status": "success",
     "sizeBytes": 734003200 }'
✓ orders-db  backup 2 h ago  ok
// what's new in 4.2

A correctness release. The numbers got honest and the roles got teeth.

Shipped between August and September 2026. The live demo runs it.

// for business leaders

Uptime numbers you can put in front of a customer

  • Measured, not estimated. Every uptime figure is built from snapshots, weighted by minutes, with the unmeasured time shown as unmeasured.
  • Early warning. An SLA group reads AT_RISK during the month, while there is still time to act.
  • Nobody has to notice an outage. A service going down opens a ticket in the right queue in under 30 seconds.
  • A status page you own. Customers and management see uptime, incidents and maintenance, no login required.
  • The data stays yours. Everything runs in your cluster. There is no vendor cloud and no per-host bill.

"What's our SLA?" gets an answer at 5-minute resolution instead of two days of Excel.

// for it leaders

Built by operators, configured from Git

  • One Helm chart. Core, UI and PostgreSQL; ingress-nginx or Istio; LDAP bind secret from your own Secret.
  • An agent worth trusting. No third-party code, no inbound port, GET-only RBAC, signed releases, MIT licence.
  • Workflows, catalogue and groups as YAML, mounted from the chart. No admin UI to drift away from what is in Git.
  • Hardened by default. Production mode forces security headers, rate limiting and introspection off. Webhooks are SSRF-checked on every redirect.
  • GraphQL and REST for everything the UI does, WebSocket subscriptions for what changes.

Monitoring tells you what is happening. ITOps is the layer above it: what it means, who owns it, and what to do.

// how itops compares

Where ITOps sits next to the tools you have probably already priced

FeatureITOpsDatadogPagerDutyUptime Robot
Self-hosted✓✗✗✗
Kubernetes-native✓✓✗✗
SLA tracking & reporting✓Partial✗Basic
Ticketing built in✓✗✗✗
Service discovery✓✓✗✗
Bare-metal support✓✓✗✗
Public status page✓✗✓✓
Licensing modelFree core + paid pluginsPer host, subscriptionPer user, subscriptionFree tier + subscription
Setup timeMinutesHoursHoursMinutes
Data ownership100% yoursVendor cloudVendor cloudVendor cloud
// pricing

Simple, transparent pricing

Self-hosted, so your data stays yours. No per-seat surprises.

Community

Free
For small teams getting started
  • Unlimited clusters & services
  • Kubernetes agent and HTTP probes
  • Operations catalogue + health dashboard
  • CMDB & dependency graph
  • Storage usage monitoring
  • Outgoing webhooks
  • Community support
  • SLA plugin
  • Ticketing plugin
Get started
recommended

Professional

Licensed
Both plugins, unlimited services
  • Everything in Community
  • SLA plugin — uptime measurement, error budgets, daily reports, public status page
  • Ticketing plugin — ITIL ticket types, GitOps workflows, service catalogue, audit trail
  • Backup monitoring
  • LDAP / Active Directory sign-in, declared in Git
  • Email support
Contact for pricing

Enterprise

Custom
Dedicated support, custom integrations
  • Everything in Professional
  • Unlimited users
  • Audit log + compliance reporting
  • On-premises deployment support
  • Named contact with its own SLA
  • Custom integrations
Contact sales
// ready?

Ready to prove your SLA?

One command to a running platform. The quickstart takes it from there to a measured service in ten minutes.

$ helm install itops itops/itops -n itops --create-namespaceclick to copy