Troubleshooting
The problems people actually hit on a fresh install, in the order they tend to hit them.
Each entry is a symptom, the cause, and the fix. Most of them come from a real installation log.
The login page loads but every login fails with a network error
Cause. The browser talks to the API cross-origin and the core only allows localhost by default.
Fix. Set env.ITOPS_SECURITY_CORS_ALLOWED_ORIGINS to the UI's origin (https://app.example.com) and make sure ui.apiUrl points at the API's public URL. With Istio the two hosts are still different origins, so the same value is needed. Confirm with the browser's network tab: a blocked preflight shows as a failed OPTIONS /graphql.
ImagePullBackOff on the SLA portal in a fresh namespace
Cause. The sla-portal chart's default imagePullSecrets references a Secret named ghcr-secret that does not exist. The image itself is public.
Fix. --set imagePullSecrets=null, or an empty list in your values file.
Agents and curl calls answer 401
Cause. The X-API-Key header does not match ITOPS_SECURITY_OPERATOR_API_KEY, or the key is empty on the server and the client sends one anyway. A POST /api/v1/backup/report with an empty server-side key logs a loud warning at startup.
Fix. Read the key back from the Secret the chart rendered and compare:
kubectl -n itops get secret itops-secrets -o jsonpath='{.data.ITOPS_SECURITY_OPERATOR_API_KEY}' | base64 -d
Every failed API-key check is also written to the SLA event log as auth_failed, with the method and path.
The agent runs but every service stays UNKNOWN
Three causes, in order of likelihood:
ITOPS_WATCHis nottrue. Without it the agent sends heartbeats only, which is what it was asked to do. Set it, and give it the Role from the deploy manifest.- The agent's
ITOPS_NODE_IDdoes not match thenodeIdthe services were registered under. The agent only asks for services on its own node. Compare the two strings exactly. - The workload has a different name than the service, or lives in another namespace. The agent looks a service up by its
nameas a Deployment, StatefulSet or DaemonSet in the service'snamespace, falling back toITOPS_NAMESPACE. A workload it cannot find is reported asUNKNOWN, deliberately notDOWN.
The agent's stdout says which of the three it is.
A pushed service went OUTAGE although the host is fine
Cause. The push stopped. External services turn UNKNOWN after two minutes without a report and OUTAGE after five. This is the platform refusing to mistake silence for health.
Fix. Check the cron line or timer on the host, then the network policy between the host and the core. MAINTENANCE is never overwritten by the stale check, so a planned stop can be announced as such.
The tree shows an unknown/… branch
Cause. A push arrived with a missing or malformed nodeId. The core pads it with the literal unknown to four segments and stores the report there rather than rejecting it, with a warning in the response body.
Fix. Read the warnings array in the response, fix the client's path, and delete the stray node with the deleteOperationsNode GraphQL mutation.
The SLA dashboard is empty on a fresh install
Cause. Either there is no licence, in which case the SLA tab is hidden and the snapshots still accumulate, or the platform is younger than its first aggregation. Snapshots are taken every five minutes and aggregated with a fifteen-minute delay so that every agent's report has arrived.
Fix. Wait twenty minutes. If licenseInfo in GraphQL says the SLA plugin is not licensed, see Licence and plugins.
The public status page shows no groups
Cause. The portal renders only the SLA groups named in its slaTargets value, or the core has never pushed a report to it.
Fix. Set slaTargets with the group names exactly as the core reports them, check ITOPS_SLA_PORTAL_URL and the matching API key on the core, and trigger a report by hand:
curl -X POST $API/api/v1/sla/report/generate -H "X-API-Key: $KEY"
The startup probe keeps failing on first boot
Cause. Migrations on a slow disk. The probe allows ten minutes.
Fix. Watch the core log; if it is still migrating, wait. If it says it cannot reach PostgreSQL, the bundled database's postgresql.auth.password and secretEnv.ITOPS_DATABASE_PASSWORD differ.
Editing groups/*.yaml does not create a group
Cause. Groups are upserted from the mounted directory at core start, not on ConfigMap change.
Fix. kubectl -n itops rollout restart deploy/itops-core after helm upgrade, or create the group in the UI.
The Auth Providers page will not let me add a provider
That page is read-only by design. Providers are declared in users/*.yaml in the chart; see Users, groups and roles.
I want to delete a node or a service
Services: the deleteService GraphQL mutation, or the delete action on the service page (admin role). Nodes: deleteOperationsNode. Both are admin-only and both are refused for the read-only viewer role.
Rate limiting can be bypassed with a header
Cause. ingress-nginx is configured with use-forwarded-headers: true and trusts X-Forwarded-For from anyone. This is a cluster setting and it affects every ingress, not only ITOps. It is described with numbers in Install.
Something else
Run the core with ITOPS_LOGGING_LEVEL=debug and ITOPS_LOGGING_FORMAT=console, reproduce, and send the relevant lines to puskas.balazs@36306800800.hu. Include the chart version and the version field from GET /health.