Kubernetes agent
itops-agentv2 - what it reports, what it will never do, the RBAC it needs and how to run it, in a cluster or as an HTTP prober anywhere.
itops-agentv2 is the component that tells the core what is actually running. It is about 1,100 lines of Go with no third-party code, built into a scratch image that contains the binary and a CA bundle and nothing else. It is small on purpose: the promise on this page is that you can read the whole agent before you give it access to your cluster, and a test in its repository reads the agent's own source and fails the build if any statement below stops being true.
Source and releases: https://github.com/balazspuskas/itops-agentv2 (MIT).
What it does
Every ITOPS_INTERVAL (default 30 s), in this order:
- Heartbeat, always.
POST /api/v1/operator/heartbeatwith its node id, version and two counters:{"nodeId":"acme/platform/prod/cluster1","version":"0.3.0","watchedServices":3,"healthyServices":2} - Watch, only with
ITOPS_WATCH=true.GET /api/v1/operator/servicesfor its own node, then oneGETper service against the Kubernetes API, looking for a Deployment, StatefulSet or DaemonSet with the service's name in its namespace. It reads replica counts and container images, nothing else, and sends onePOST /api/v1/operator/status:The rules: every replica ready is{"name":"payment-api","externalId":"acme/platform/prod/cluster1/payment-api", "namespace":"prod","status":"DEGRADED","message":"2/3 replicas ready", "workloadType":"Deployment","workloadName":"payment-api", "replicas":3,"readyReplicas":2,"availableReplicas":2, "images":["ghcr.io/acme/payment-api:1.2.3"]}OPERATIONAL; some ready isDEGRADED; none ready, or scaled to zero, isDOWN; a workload it cannot find isUNKNOWN, deliberately notDOWN, because "not found" and "crashed" are different facts. - Probe, only with
ITOPS_TARGETSset. One timedGETper target, posted toPOST /api/v1/health/report. See HTTP probes.
The agent has no discovery of its own. It does not read ConfigMaps or labels; it looks for exactly the names the core hands it, which is why registration comes first.
What it will never do
- Listen. It never calls
net.Listen. No port, no metrics endpoint, no admin socket. - Talk to anyone else.
ITOPS_URLalways; the Kubernetes API server only when watching; theITOPS_TARGETSendpoints only when set. No telemetry, no update check. - Write to Kubernetes. Every request its Kubernetes client can build is a
GET, and a test asserts that on the wire. - Run anything. It never imports
os/exec; there is no shell in the image anyway. - Write to disk. No state file, no cache, no lock file. Log lines on stdout are the only output.
- Collect from the host. No hostnames, node IPs, process lists, environment dumps, logs, Secrets or ConfigMap contents. Images are reported because a version is what you want to know when a deployment starts failing.
- Obey a probe target. Probe responses are drained and discarded, never parsed.
- Act on server replies. The service list is decoded into a handful of string fields; everything else in the response is ignored.
With ITOPS_WATCH=true it reads three files, all under /var/run/secrets/kubernetes.io/serviceaccount/: the token, the cluster CA and the namespace. With watching off it reads no files at all.
Configuration
| Variable | Required | Default | Meaning |
|---|---|---|---|
ITOPS_URL |
yes | Base URL of the core, http:// or https://. Inside the cluster: http://itops-core.itops.svc:8080. |
|
ITOPS_API_KEY |
yes | The operator API key. Sent only to ITOPS_URL, only as X-API-Key. |
|
ITOPS_NODE_ID |
yes | organization/platform/environment/cluster. Must match the nodeId the services were registered under. |
|
ITOPS_INTERVAL |
no | 30s |
Cycle length. 45s, 2m, or a bare number of seconds. |
ITOPS_TIMEOUT |
no | 10s |
Per-request timeout, to either server. Must be shorter than the interval. |
ITOPS_WATCH |
no | false |
Look workloads up in Kubernetes and report them. Needs the Role below. |
ITOPS_NAMESPACE |
no | own namespace | Where to look when the service has no namespace of its own. |
ITOPS_TARGETS |
no | HTTP endpoints to probe. Needs no cluster at all. |
Bad configuration exits with status 2 before sending anything. Asking to watch while not running in a cluster is fatal too: silently degrading to heartbeat-only while someone believes their services are watched is the failure mode this refuses.
Deploy in a cluster
The manifest in the repository is complete and short enough to read in full: a ServiceAccount, a Role, a RoleBinding and a Deployment. Edit the three values and apply it in the namespace whose workloads you want watched.
kubectl -n shop create secret generic itops-agentv2 --from-literal=api-key='<operator API key>'
curl -sO https://raw.githubusercontent.com/balazspuskas/itops-agentv2/main/deploy/kubernetes.yaml
kubectl -n shop apply -f kubernetes.yaml
The parts that matter:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role # a Role, per namespace, not a ClusterRole
rules:
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets", "daemonsets"]
verbs: ["get"] # no list, no watch, no write
---
kind: Deployment
spec:
template:
spec:
serviceAccountName: itops-agentv2
containers:
- name: agent
image: ghcr.io/balazspuskas/itops-agentv2:latest
env:
- { name: ITOPS_URL, value: http://itops-core.itops.svc:8080 }
- { name: ITOPS_NODE_ID, value: acme/shop/prod/eu-1 }
- { name: ITOPS_WATCH, value: "true" }
- name: ITOPS_API_KEY
valueFrom: { secretKeyRef: { name: itops-agentv2, key: api-key } }
resources:
requests: { cpu: 5m, memory: 16Mi }
limits: { memory: 32Mi }
securityContext:
runAsNonRoot: true
runAsUser: 65532
readOnlyRootFilesystem: true
allowPrivilegeEscalation: false
capabilities: { drop: [ALL] }
get without list is deliberate: the agent can look up names the core gave it and cannot enumerate what else lives in the namespace. Watching several namespaces means copying the Role and RoleBinding into each one. That is more typing than a ClusterRole, and it makes a cluster-wide read grant a decision somebody took rather than a default.
Without ITOPS_WATCH the agent needs no ServiceAccount permissions at all, and automountServiceAccountToken: false is appropriate.
HTTP probes
The same binary watches things that are not in Kubernetes: a managed database's health endpoint, an appliance, a SaaS status URL, a VM's /healthz. ITOPS_TARGETS is a list of key=value records, one per line or separated by ;:
name=payments,url=https://pay.internal/healthz,criticality=high,slaGroup=checkout-flow
name=db,url=http://db.internal:8080/health,status=200,slaGroup=orders-database
name=cdn,url=https://www.example.com/,displayName=CDN edge,type=external
name and url are required. status names the exact HTTP code that counts as healthy; without it any 2xx does. criticality, slaGroup, type and displayName are recorded when the service is first created. An unknown key or a malformed record is fatal at startup, not a silently skipped check.
Each cycle sends one report per target:
{"nodeId":"acme/platform/prod/cluster1","service":"payments","status":"OPERATIONAL",
"message":"HTTP 200 in 12ms","criticality":"high","tags":["http-probe","agentv2"]}
A timeout, a refused connection or a wrong status code is DOWN with the reason in message. The probes need no cluster, so this mode also runs as a plain container on a VM:
docker run --rm \
-e ITOPS_URL=https://api.example.com -e ITOPS_API_KEY=… \
-e ITOPS_NODE_ID=acme/platform/prod/dc-budapest \
-e ITOPS_TARGETS='name=galera,url=http://10.0.0.5:9200/health' \
ghcr.io/balazspuskas/itops-agentv2:latest
Image and supply chain
FROM scratch, uid 65532, no shell, no package manager, no libc. Releases are built by GitHub Actions and signed keyless with Sigstore, with an SBOM and SLSA provenance attached:
cosign verify \
--certificate-identity-regexp '^https://github.com/balazspuskas/itops-agentv2/\.github/workflows/release\.yml@refs/' \
--certificate-oidc-issuer 'https://token.actions.githubusercontent.com' \
ghcr.io/balazspuskas/itops-agentv2:<version>
Or build it yourself: go build ./cmd/agentv2 has no dependencies to fetch.
Reading the agent's output
Every cycle logs one line per outcome. The three questions it answers are the three reasons a service stays UNKNOWN: watching is off, the node id does not match the registration, or the workload is not where the service says it is. Troubleshooting walks through them.
The old operator
Versions before 4.2 shipped a Kubernetes operator that discovered services from labelled ConfigMaps. It is retired. Its chart, itops/itops-agent 1.4.1, is still on the chart repository for installations that have not migrated, and the ConfigMap it read is no longer needed: register the service once and let agentv2 report on it.