An IT operations platform that runs on your own cluster, for your Kubernetes workloads and for everything running next to them. One Helm install, and uptime numbers your board can read: measured, not estimated.
The demo runs Meridian Retail Group, a fictional online retailer: a Kubernetes commerce platform, a bare-metal database cluster and a few cloud services. Sign in with demo / MeridianDemo2026. The account reads everything and changes nothing.
$ helm repo add itops https://charts.mlops.hu $ helm install itops itops/itops -n itops --create-namespace NAME: itops · STATUS: deployed · REVISION: 1 $ curl -X POST $API/api/v1/operator/register \ -H "X-API-Key: $KEY" -d '{ "nodeId": "meridian/commerce/prod/eu-central-1", "name": "payment-gateway", "criticality": "critical", "slaGroup": "checkout-flow", "autoIncident": true }' {"success":true,"created":true} # 30 seconds later the agent reports in ✓ payment-gateway OPERATIONAL 3/3 ready 99.99% budget 92% $
The public status page of the demo estate, as the SLA plugin publishes it: 90 days per group, the month's uptime against its target, the error budget left. When the platform was blind, the day is grey, not green.
One self-hosted platform. The core is free; the SLA and ticketing plugins unlock with a licence key.
Uptime is built from 5-minute snapshots on an exact-duration timeline and weighted by minutes, so a large service moves the headline number and a small one does not. Maintenance windows are excluded from both sides of the fraction. Time the platform could not measure is reported as UNKNOWN, never as a silent 100%.
Every SLA group shows its budget in minutes: allowed, used, remaining, and the burn rate. A month in flight is flagged AT_RISK before the budget is spent, not after the report is printed. Four seeded tiers from 99.0% to 99.99%, with your own targets per group.
Incidents, service requests, problems and changes, kept apart so MTTR and request fulfilment stop blurring into one number. Priority is derived from impact × urgency, and the priority sets the deadlines: a P1 is answered in 15 minutes and resolved in 4 hours. Putting a ticket on hold stops the clock.
When a service goes DOWN, the platform opens an SLA incident and a ticket in the right queue, with type, impact and urgency filled in. When the service comes back and nobody has touched the ticket, it resolves itself and says why.
itops-agentv2 is about 1,100 lines of Go with no third-party code, shipped in a scratch image. It reports replica counts and images every 30 seconds, needs get on workloads and nothing else, opens no port and never writes to your cluster. MIT licensed.
The same agent probes HTTP endpoints with no cluster at all. A 200-line POSIX shell agent covers hosts that only have curl. Anything that can send one HTTP request lands on the same dashboard with the same SLA maths, and a service that stops reporting turns UNKNOWN after 2 minutes and OUTAGE after 5.
A standalone page for customers and management: uptime per SLA group, error budget left, incident history and maintenance windows for the whole month. No login, just a link. When the data is thin, it says so instead of showing green.
Every service carries its team, owner, on-call, runbook and dashboard links, tags and tier. Dependencies are declared once and rendered in both directions, so "what breaks if this goes down" is one query, not a meeting.
Any backup tool ends with one HTTP call, and an overdue backup raises an alert. Storage reports turn into healthy / warning / critical at 30% and 10% free. 25 event types go out as HMAC-signed webhooks with SSRF protection on every hop. Four roles from group membership, a read-only viewer role, field-level permissions, LDAP and Active Directory sign-in.
The catalogue belongs to the server. You tell it what should exist; the agents tell it what does. Both are plain HTTP calls you can put in a Helm hook, a CI job or a cron line.
Core, UI and a bundled PostgreSQL from one chart. Bring your own database, ingress and LDAP when you are ready for production.
$ helm repo add itops https://charts.mlops.hu $ helm install itops itops/itops \ -n itops --create-namespace first login: admin / Password123! change it. it is on the checklist.
Every service is identified by one path: org/platform/env/cluster/service. The name and the node are the only required fields. Everything else unlocks one more feature at a time.
$ curl -X POST $API/api/v1/operator/register \
-H "X-API-Key: $KEY" -d '{
"nodeId": "acme/shop/prod/eu-1",
"name": "orders-db",
"slaGroup": "orders-database",
"operations": { "backup":
{ "expected": true, "maxAgeDays": 1 } }
}'The Kubernetes agent looks each registered service up and reports its state. Hosts push their own health. Backup jobs end with one more request.
# any backup script, last line $ curl -X POST $API/api/v1/backup/report \ -H "X-API-Key: $KEY" -d '{ "path": "acme/shop/prod/eu-1/orders-db", "status": "success", "sizeBytes": 734003200 }' ✓ orders-db backup 2 h ago ok
Shipped between August and September 2026. The live demo runs it.
AT_RISK projection for the month in flight.ON_HOLD pauses the clock.itops-agentv2 is stdlib-only, read-only, and small enough to audit in an afternoon.AT_RISK during the month, while there is still time to act."What's our SLA?" gets an answer at 5-minute resolution instead of two days of Excel.
Monitoring tells you what is happening. ITOps is the layer above it: what it means, who owns it, and what to do.
| Feature | ITOps | Datadog | PagerDuty | Uptime Robot |
|---|---|---|---|---|
| Self-hosted | ✓ | ✗ | ✗ | ✗ |
| Kubernetes-native | ✓ | ✓ | ✗ | ✗ |
| SLA tracking & reporting | ✓ | Partial | ✗ | Basic |
| Ticketing built in | ✓ | ✗ | ✗ | ✗ |
| Service discovery | ✓ | ✓ | ✗ | ✗ |
| Bare-metal support | ✓ | ✓ | ✗ | ✗ |
| Public status page | ✓ | ✗ | ✓ | ✓ |
| Licensing model | Free core + paid plugins | Per host, subscription | Per user, subscription | Free tier + subscription |
| Setup time | Minutes | Hours | Hours | Minutes |
| Data ownership | 100% yours | Vendor cloud | Vendor cloud | Vendor cloud |
Self-hosted, so your data stays yours. No per-seat surprises.
One command to a running platform. The quickstart takes it from there to a measured service in ten minutes.
$ helm install itops itops/itops -n itops --create-namespaceclick to copy