Operations

Troubleshooting

Common problems, how to confirm the cause, and how to fix them.

Most checks below need the discovery leader. Find it with:

kubectl exec -n telark telark-redis-master-0 -- redis-cli GET election:prewarm

If your problem is not here, open an issue with the symptom and the relevant kubectl logs and kubectl get output.

The chart refuses to render

Error mentionsFix
bootstrap adminsAdd --set 'app.auth.bootstrap.admins={you@example.com}'.
storage class, ReadWriteManySet app.persistence.storageClass to a ReadWriteMany class, or use --set app.singleNode=true on a single node.
passkey id or originWith an Ingress or Gateway, set app.auth.passkey.id and app.auth.passkey.origin.
failOpenapp.kyverno.failOpen and kyverno.features.forceFailurePolicyIgnore.enabled must match.

Passkey registration or sign-in fails

Confirm: the dashboard shows a warning about a secure origin, or the browser reports a WebAuthn error.

Fix:

  • Serve the dashboard over HTTPS with a certificate the browser trusts. Only https:// and http://localhost allow passkeys.
  • Set app.auth.passkey.id to the hostname (telark.example.com) and app.auth.passkey.origin to the full origin (https://telark.example.com), then helm upgrade.
  • A passkey registered on one hostname does not work on another. Register again on the new host, or use Settings → Security → Passkeys → Add on another device.

Enrolment tokens last 10 minutes. Run break-glass --enroll again for a new one.

Applications page stays empty

Confirm:

  • The namespace is not excluded: check Settings → Governance → Excluded namespaces. default and the kube-* namespaces are excluded by default.
  • The workloads carry a grouping label such as app.kubernetes.io/name or app.
  • Discovery synced its informers: kubectl logs -n telark deploy/telark-discovery-service | grep 'informers synced'.
  • GET /api/v1/discovery/status shows a pass that finishes.

Fix: Use a non-excluded namespace or remove it from the list. With a custom install, make sure the discovery ServiceAccount has the chart's cluster-wide read permissions.

Plan stuck in scheduled

Confirm: the start time has passed, and the discovery leader logs This replica is now leader without repeated leadership losses.

Fix: Wait one controller tick (31 seconds by default). If the phase still does not move, delete the leader pod. Another replica takes over within about 15 seconds.

Plan health is degraded or drifted

Confirm: open the plan's Health panel. Each row says which policy is missing, not ready, or different from what the plan expects. List the deployed policies:

kubectl get policies.kyverno.io -A -l telark.io/protection-plan=<plan-id>

Fix:

  • Drifted: someone edited a policy. Telark redeploys it on the next tick. If drift keeps coming back, find what in the cluster keeps editing it.
  • Degraded: a policy is missing or not ready. Check Kyverno: kubectl get pods -n telark | grep kyverno. Telark redeploys missing policies on the next tick once Kyverno is healthy.
  • Refresh health on the plan page recomputes the status now; repairs still happen on the tick.

An enforce plan did not block a change

Confirm: Kyverno's admission pods were ready at the time, and the resource is in the plan's scope and not excluded.

Fix: With the default app.kyverno.failOpen=true, requests are admitted while Kyverno is down or rolling out. For strict enforcement, see Install Telark for production.

A finished plan shows no violations

This is expected. When a plan ends, Telark deletes its policies and its live violation list is empty. The history is in the plan's final report under Reports.

Rollback stuck in pending

Confirm: the leader's logs: kubectl logs -n telark <leader-pod> | grep -i rollback. picked up pending entry should appear within seconds. lock for <app> held elsewhere; deferring pickup means another rollback of the same application is still running.

Fix: Wait for the other rollback. If nothing is picked up, restart the discovery leader.

Rollback stuck in in_progress or failed

Confirm: the rollback's error field, and kubectl describe on the resource it names. Common causes are an admission rejection (often an enforce protection plan on the same scope), an immutable field, or a missing namespace.

Fix: Remove the cause and roll back again. A rollback still in_progress after 5 minutes is marked failed.

Snapshot volume is full

Confirm: Settings → Governance → Snapshot storage is near 100%, and the exporter logs failed to write snapshot to disk:

kubectl logs -n telark deploy/telark-exporter-service | grep -i snapshot

Fix:

  • Now: lower the snapshots kept per application. Each application is pruned at its next change. See Change snapshot retention.
  • Next: expand telark-exporter-snapshots-pvc if your storage class allows expansion.
  • Later: exclude noisy namespaces.

Google sign-in fails

See the troubleshooting table in Set up Google sign-in.

Notifications stop appearing

Confirm: the exporter stores in-app notifications in Redis. kubectl logs -n telark deploy/telark-exporter-service | grep -i redis.

Fix: Make sure telark-redis-master-0 is running, then restart the exporter.

Insights show rule text without an explanation

Confirm: the model has not been downloaded yet, or the install is air-gapped.

Fix: In connected installs, wait for the download or choose Install model in Settings → Local analyzer. Air-gapped installs must pre-load the model.

CRDs or the namespace stuck in Terminating

Users, groups and access roles carry cleanup finalizers that only the auth service removes. Clear them with step 1 of Uninstall Telark → Full teardown; the deletions then complete.