Operations

Health checks

The liveness and readiness routes each Telark service exposes, and how to check the whole install.

Check the whole install

kubectl get pods -n telark
helm test telark -n telark

helm test starts a pod that calls the readiness route of every service that has one, then checks that the Kubernetes API serves the telark.io group. It should end with Phase: Succeeded.

Probe routes

Every API service serves both routes on its http port (8080). They need no authentication.

ProbeRoute
LivenessGET /api/v1/status/live
ReadinessGET /api/v1/status/ready

What "ready" means per service:

ServiceReady when
exporterRedis is reachable and its application cache has synced
discoveryRedis is reachable. Startup, exporter and Kubernetes problems show as degraded, not unready
authThe process is up
notifierConnected to NATS
analyzerRedis answers a ping
uiNo probes. The dashboard is static files behind nginx and serves only GET /healthz, so the chart sets services.ui.includeHealthCheck: false

To call a route yourself:

kubectl port-forward -n telark svc/telark-discovery-service 8080:8080
curl -s localhost:8080/api/v1/status/ready

Probe timings

All services share one block, app.shared.healthCheck:

SettingLivenessReadiness
initialDelaySeconds155
periodSeconds155
timeoutSeconds1515
failureThreshold33

The defaults allow for slow cold starts. Lower them only after you have measured pod start times in your cluster.

Next