VMs and bare metal
Three ways to get a machine without Kubernetes onto the dashboard - the shell agent, the HTTP prober, and a curl in cron.
Anything that can send one HTTP request can be a service in ITOps, with the same status model, the same SLA maths and the same tickets as a Kubernetes workload. There are three ways in, from the least to the most informative.
1. A heartbeat from the host: itops-agent-sh
For machines where installing a binary is more trouble than it is worth, itops-agent-sh does the agent's heartbeat in about 200 lines of POSIX shell, half of them comments. It needs a shell and curl, is tested under dash, and uses no jq, no Python and no bashisms.
Source: https://github.com/balazspuskas/itops-agent-sh (MIT).
What it sends, and the only thing it sends:
POST $ITOPS_URL/api/v1/operator/heartbeat
{"nodeId":"acme/platform/prod/dc-budapest","version":"0.1.0","watchedServices":0,"healthyServices":0}
Nothing is collected from the machine: no hostname, IP, uname, process list or disk usage. The only fact that leaves the host is the node id you typed in. The node shows up in the tree with its last-seen time; services under it come from the other two methods.
| Variable | Required | Default | Meaning |
|---|---|---|---|
ITOPS_URL |
yes | Base URL, http:// or https://. |
|
ITOPS_API_KEY_FILE |
one of the two | Read the key from this file. Preferred. | |
ITOPS_API_KEY |
The key in the environment. | ||
ITOPS_NODE_ID |
yes | organization/platform/environment/cluster. Validated: letters, digits, ., -, _, /. |
|
ITOPS_INTERVAL |
no | 30 |
Seconds between heartbeats in --loop mode. Plain integer. |
ITOPS_TIMEOUT |
no | 10 |
Seconds before a request gives up. |
Prefer the file: an environment variable is readable from /proc/<pid>/environ, and a key in a systemd unit shows up in systemctl show, the journal and your configuration management. A root-owned 0640 file avoids all three. The key never appears on curl's command line either; it is passed on stdin, because argv is visible in ps to every user on the box.
Install as a systemd timer, which retries a missed run and keeps nothing resident between heartbeats:
install -m 0755 itops-agent.sh /usr/local/bin/itops-agent.sh
groupadd --system itops
install -d -m 0750 -g itops /etc/itops
install -m 0640 -g itops /dev/null /etc/itops/api-key
printf '%s' 'YOUR-KEY' > /etc/itops/api-key
install -m 0644 systemd/itops-agent.service systemd/itops-agent.timer /etc/systemd/system/
# edit ITOPS_URL and ITOPS_NODE_ID in the .service unit
systemctl daemon-reload
systemctl enable --now itops-agent.timer
systemctl list-timers itops-agent.timer
The shipped unit runs every 30 seconds with a 5-second random delay so cloned machines do not stampede the server on the same tick, as DynamicUser with an empty capability set, ProtectSystem=strict, MemoryDenyWriteExecute and a @system-service syscall filter. A cron line works too:
* * * * * ITOPS_URL=https://api.example.com ITOPS_NODE_ID=acme/platform/prod/dc-budapest ITOPS_API_KEY_FILE=/etc/itops/api-key /usr/local/bin/itops-agent.sh --once
Exit codes: 0 accepted, 1 the request failed, 2 the configuration was unusable and nothing was sent.
2. Probe it from anywhere: itops-agentv2 with ITOPS_TARGETS
If the machine exposes a health URL, you do not need anything on it. Run the Kubernetes agent, in a cluster or as a plain container, with a target list, and it reports each URL's status every cycle. This is the right tool for managed databases, appliances, SaaS endpoints and anything with a /healthz:
ITOPS_TARGETS='name=galera-node-1,url=http://10.0.0.5:9200/health,slaGroup=orders-database,criticality=critical
name=wms,url=https://wms.internal/ping,status=200,displayName=Warehouse WMS'
Details on the agent page.
3. Let the host say how it feels: curl in cron
The richest option is a script on the host that knows what "healthy" means there and pushes a status with a message. This is a Galera node reporting its own cluster state every minute:
#!/bin/sh
# /etc/cron.d/itops-health: * * * * * root /usr/local/bin/itops-health.sh
size=$(mysql -Nse "SHOW STATUS LIKE 'wsrep_cluster_size'" | cut -f2)
ready=$(mysql -Nse "SHOW STATUS LIKE 'wsrep_ready'" | cut -f2)
if [ "$ready" = "ON" ] && [ "$size" -ge 3 ]; then s=OPERATIONAL; elif [ "$ready" = "ON" ]; then s=DEGRADED; else s=DOWN; fi
curl -sS -X POST "$ITOPS_URL/api/v1/health/report" \
-H "X-API-Key: $(cat /etc/itops/api-key)" -H "Content-Type: application/json" \
-d "{\"path\":\"acme/platform/prod/dc-budapest/galera-node-1\",\"status\":\"$s\",\"message\":\"wsrep_cluster_size=$size, wsrep_ready=$ready\",\"criticality\":\"critical\",\"slaGroup\":\"orders-database\"}"
The first push creates the service. Subsequent pushes update it. Disk usage and backup outcomes go to the other two endpoints in the same style; all three are on the push API page.
What happens when the host goes quiet
A pushed service that stops reporting turns UNKNOWN after two minutes and OUTAGE after five. This is the platform refusing to mistake silence for health: a dead host looks dead, and the SLA counts it. A planned stop should be pushed as MAINTENANCE, which the stale check never overwrites, or covered by an exclusion window.
Which one
| You have | Use |
|---|---|
| a health URL reachable from somewhere the agent can run | HTTP probes, nothing on the host |
a host with curl and a way to know if it is healthy |
curl in cron, with a real message |
| a host you only want to see alive in the tree | itops-agent-sh |
They combine. The demo estate's Budapest datacenter uses all three.