Public status page
The SLA portal - a standalone status page fed by the daily report, with nothing else on it, safe to put on a public hostname.
The SLA portal is a small Go service with its own SQLite database. The core pushes the daily report to it; it renders a status page. It has no access to the platform, no login and no write API beyond the ingest endpoint, which is why it can live on a public hostname while the platform stays inside.
What visitors see, per published SLA group: uptime for the month against the target, error budget remaining, incident history with durations, maintenance windows, a 30-day trend and a projection of when the budget runs out at the current burn rate. When a group has too little data, the page says so instead of showing green.
Install
helm install sla-portal itops/sla-portal --version 1.4.0 -n sla-portal --create-namespace \
--set imagePullSecrets=null \
--set ingress.host=status.example.com \
--set apiKey="$(openssl rand -hex 24)" \
-f targets.yaml
| Value | Default | Notes |
|---|---|---|
image.repository / image.tag |
ghcr.io/mlops-app/sla-portal / 1.4.0 |
public image |
imagePullSecrets |
[{name: ghcr-secret}] |
set to null; the default names a Secret that does not exist in a fresh namespace |
apiKey |
empty | required in practice; empty leaves the ingest endpoint unauthenticated |
ingress.host |
sla.mlops.hu |
your hostname |
ingress.tls.enabled / ingress.tls.secretName |
true / sla-portal-tls |
the ingress asks cert-manager's letsencrypt-prod issuer; the issuer name is not configurable in 1.4.0 |
persistence.size / persistence.storageClass |
1Gi / cluster default |
the SQLite volume |
slaTargets |
{} |
what to publish, see below |
resources |
50 m / 64 Mi requests, 200 m / 128 Mi limits |
What to publish
Only groups listed in slaTargets appear. Keys are the group names exactly as the core reports them; uptime is the target shown next to the measured figure; label is the public name.
# targets.yaml
slaTargets:
checkout-flow:
uptime: 99.95
label: "Online checkout"
orders-database:
uptime: 99.99
label: "Order database"
back-office:
uptime: 99.5
label: "Warehouse and back office"
An internal group you would rather not publish is simply left out.
Wire the core to it
The core needs the portal's in-cluster address and the same key; the UI needs the public address for its "SLA portal" link.
# itops values
env:
ITOPS_SLA_PORTAL_URL: http://sla-portal.sla-portal.svc.cluster.local
ITOPS_SLA_PORTAL_API_KEY: "<the same key>"
ui:
slaPortalUrl: https://status.example.com
networkPolicy:
allowedEgressNamespaces: [sla-portal]
From then on every daily report is pushed as POST /api/v1/report with Authorization: Bearer <key>. Trigger one now rather than waiting for midnight:
curl -X POST $API/api/v1/sla/report/generate -H "X-API-Key: $KEY"
The portal's own API
Everything the page shows is available as JSON, unauthenticated, for your own dashboards:
| Endpoint | Returns |
|---|---|
GET /api/v1/summary |
current month per published group: uptime, target, budget, status, projection |
GET /api/v1/services |
published groups and their services |
GET /api/v1/services/{id}/history?days=30 |
daily history, up to 365 days |
GET /api/v1/reports |
the reports received |
GET /api/v1/reports/{date} |
one day's report in full |
GET /api/v1/targets |
the slaTargets in effect |
GET /health |
liveness |
The page itself is static HTML served from /, so it is easy to put behind a CDN.
Running it elsewhere
The portal is one container with one volume and five environment variables, so it does not have to be in the cluster. Anywhere the core can reach over HTTP works:
| Variable | Default |
|---|---|
PORT |
8080 |
DB_PATH |
/data/sla-portal.db |
API_KEY |
none; set it |
SLA_TARGETS |
JSON with the same shape as slaTargets |
What it does not do
It does not query the platform, cannot open tickets, holds no user data and has no admin interface. If it is compromised, the attacker learns your uptime.