docs / Feeding data in

Kubernetes agent

itops-agentv2 - what it reports, what it will never do, the RBAC it needs and how to run it, in a cluster or as an HTTP prober anywhere.

itops-agentv2 is the component that tells the core what is actually running. It is about 1,100 lines of Go with no third-party code, built into a scratch image that contains the binary and a CA bundle and nothing else. It is small on purpose: the promise on this page is that you can read the whole agent before you give it access to your cluster, and a test in its repository reads the agent's own source and fails the build if any statement below stops being true.

Source and releases: https://github.com/balazspuskas/itops-agentv2 (MIT).

What it does

Every ITOPS_INTERVAL (default 30 s), in this order:

  1. Heartbeat, always. POST /api/v1/operator/heartbeat with its node id, version and two counters:
    {"nodeId":"acme/platform/prod/cluster1","version":"0.3.0","watchedServices":3,"healthyServices":2}
    
  2. Watch, only with ITOPS_WATCH=true. GET /api/v1/operator/services for its own node, then one GET per service against the Kubernetes API, looking for a Deployment, StatefulSet or DaemonSet with the service's name in its namespace. It reads replica counts and container images, nothing else, and sends one POST /api/v1/operator/status:
    {"name":"payment-api","externalId":"acme/platform/prod/cluster1/payment-api",
     "namespace":"prod","status":"DEGRADED","message":"2/3 replicas ready",
     "workloadType":"Deployment","workloadName":"payment-api",
     "replicas":3,"readyReplicas":2,"availableReplicas":2,
     "images":["ghcr.io/acme/payment-api:1.2.3"]}
    
    The rules: every replica ready is OPERATIONAL; some ready is DEGRADED; none ready, or scaled to zero, is DOWN; a workload it cannot find is UNKNOWN, deliberately not DOWN, because "not found" and "crashed" are different facts.
  3. Probe, only with ITOPS_TARGETS set. One timed GET per target, posted to POST /api/v1/health/report. See HTTP probes.

The agent has no discovery of its own. It does not read ConfigMaps or labels; it looks for exactly the names the core hands it, which is why registration comes first.

What it will never do

  • Listen. It never calls net.Listen. No port, no metrics endpoint, no admin socket.
  • Talk to anyone else. ITOPS_URL always; the Kubernetes API server only when watching; the ITOPS_TARGETS endpoints only when set. No telemetry, no update check.
  • Write to Kubernetes. Every request its Kubernetes client can build is a GET, and a test asserts that on the wire.
  • Run anything. It never imports os/exec; there is no shell in the image anyway.
  • Write to disk. No state file, no cache, no lock file. Log lines on stdout are the only output.
  • Collect from the host. No hostnames, node IPs, process lists, environment dumps, logs, Secrets or ConfigMap contents. Images are reported because a version is what you want to know when a deployment starts failing.
  • Obey a probe target. Probe responses are drained and discarded, never parsed.
  • Act on server replies. The service list is decoded into a handful of string fields; everything else in the response is ignored.

With ITOPS_WATCH=true it reads three files, all under /var/run/secrets/kubernetes.io/serviceaccount/: the token, the cluster CA and the namespace. With watching off it reads no files at all.

Configuration

Variable Required Default Meaning
ITOPS_URL yes Base URL of the core, http:// or https://. Inside the cluster: http://itops-core.itops.svc:8080.
ITOPS_API_KEY yes The operator API key. Sent only to ITOPS_URL, only as X-API-Key.
ITOPS_NODE_ID yes organization/platform/environment/cluster. Must match the nodeId the services were registered under.
ITOPS_INTERVAL no 30s Cycle length. 45s, 2m, or a bare number of seconds.
ITOPS_TIMEOUT no 10s Per-request timeout, to either server. Must be shorter than the interval.
ITOPS_WATCH no false Look workloads up in Kubernetes and report them. Needs the Role below.
ITOPS_NAMESPACE no own namespace Where to look when the service has no namespace of its own.
ITOPS_TARGETS no HTTP endpoints to probe. Needs no cluster at all.

Bad configuration exits with status 2 before sending anything. Asking to watch while not running in a cluster is fatal too: silently degrading to heartbeat-only while someone believes their services are watched is the failure mode this refuses.

Deploy in a cluster

The manifest in the repository is complete and short enough to read in full: a ServiceAccount, a Role, a RoleBinding and a Deployment. Edit the three values and apply it in the namespace whose workloads you want watched.

kubectl -n shop create secret generic itops-agentv2 --from-literal=api-key='<operator API key>'
curl -sO https://raw.githubusercontent.com/balazspuskas/itops-agentv2/main/deploy/kubernetes.yaml
kubectl -n shop apply -f kubernetes.yaml

The parts that matter:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role                          # a Role, per namespace, not a ClusterRole
rules:
  - apiGroups: ["apps"]
    resources: ["deployments", "statefulsets", "daemonsets"]
    verbs: ["get"]                  # no list, no watch, no write
---
kind: Deployment
spec:
  template:
    spec:
      serviceAccountName: itops-agentv2
      containers:
        - name: agent
          image: ghcr.io/balazspuskas/itops-agentv2:latest
          env:
            - { name: ITOPS_URL,     value: http://itops-core.itops.svc:8080 }
            - { name: ITOPS_NODE_ID, value: acme/shop/prod/eu-1 }
            - { name: ITOPS_WATCH,   value: "true" }
            - name: ITOPS_API_KEY
              valueFrom: { secretKeyRef: { name: itops-agentv2, key: api-key } }
          resources:
            requests: { cpu: 5m, memory: 16Mi }
            limits: { memory: 32Mi }
          securityContext:
            runAsNonRoot: true
            runAsUser: 65532
            readOnlyRootFilesystem: true
            allowPrivilegeEscalation: false
            capabilities: { drop: [ALL] }

get without list is deliberate: the agent can look up names the core gave it and cannot enumerate what else lives in the namespace. Watching several namespaces means copying the Role and RoleBinding into each one. That is more typing than a ClusterRole, and it makes a cluster-wide read grant a decision somebody took rather than a default.

Without ITOPS_WATCH the agent needs no ServiceAccount permissions at all, and automountServiceAccountToken: false is appropriate.

HTTP probes

The same binary watches things that are not in Kubernetes: a managed database's health endpoint, an appliance, a SaaS status URL, a VM's /healthz. ITOPS_TARGETS is a list of key=value records, one per line or separated by ;:

name=payments,url=https://pay.internal/healthz,criticality=high,slaGroup=checkout-flow
name=db,url=http://db.internal:8080/health,status=200,slaGroup=orders-database
name=cdn,url=https://www.example.com/,displayName=CDN edge,type=external

name and url are required. status names the exact HTTP code that counts as healthy; without it any 2xx does. criticality, slaGroup, type and displayName are recorded when the service is first created. An unknown key or a malformed record is fatal at startup, not a silently skipped check.

Each cycle sends one report per target:

{"nodeId":"acme/platform/prod/cluster1","service":"payments","status":"OPERATIONAL",
 "message":"HTTP 200 in 12ms","criticality":"high","tags":["http-probe","agentv2"]}

A timeout, a refused connection or a wrong status code is DOWN with the reason in message. The probes need no cluster, so this mode also runs as a plain container on a VM:

docker run --rm \
  -e ITOPS_URL=https://api.example.com -e ITOPS_API_KEY=… \
  -e ITOPS_NODE_ID=acme/platform/prod/dc-budapest \
  -e ITOPS_TARGETS='name=galera,url=http://10.0.0.5:9200/health' \
  ghcr.io/balazspuskas/itops-agentv2:latest

Image and supply chain

FROM scratch, uid 65532, no shell, no package manager, no libc. Releases are built by GitHub Actions and signed keyless with Sigstore, with an SBOM and SLSA provenance attached:

cosign verify \
  --certificate-identity-regexp '^https://github.com/balazspuskas/itops-agentv2/\.github/workflows/release\.yml@refs/' \
  --certificate-oidc-issuer 'https://token.actions.githubusercontent.com' \
  ghcr.io/balazspuskas/itops-agentv2:<version>

Or build it yourself: go build ./cmd/agentv2 has no dependencies to fetch.

Reading the agent's output

Every cycle logs one line per outcome. The three questions it answers are the three reasons a service stays UNKNOWN: watching is off, the node id does not match the registration, or the workload is not where the service says it is. Troubleshooting walks through them.

The old operator

Versions before 4.2 shipped a Kubernetes operator that discovered services from labelled ConfigMaps. It is retired. Its chart, itops/itops-agent 1.4.1, is still on the chart repository for installations that have not migrated, and the ConfigMap it read is no longer needed: register the service once and let agentv2 report on it.

Documentation for ITOps 4.2 · charts itops 2.0.0, sla-portal 1.4.0 · rendered 2026-09-11