Operations

Metrics and observability

What Telark exposes for monitoring today, where its state is recorded, and what to alert on.

Telark does not export Prometheus metrics yet: no service serves /metrics. You observe it through probes, logs, the state it records on its custom resources, and a few read endpoints.

Where the state lives

SourceWhat it tells you
kubectl get tapp -n telark -o yamlPer application: status.health, status.lastForceSync (phase and error of the last force sync), status.rollbacks[] with each outcome, status.snapshots[], and the Published condition.
kubectl get tplan -n telark -o yamlPer protection plan: status.phase, status.health, the per-policy status.healthDetail, status.approval and status.renderedPolicies.
GET /api/v1/discovery/status on discoveryThe current discovery pass: inProgress, enqueued, remaining, startedAt, finishedAt, intervalSeconds. Needs the ReadOnly role on applications.
GET /api/v1/snapshots on the exporterSnapshot volume capacity, usage and snapshot count. The dashboard shows it under Settings → Governance → Snapshot storage.

Use the fully qualified names (applications.telark.io) or the short names (tapp, tplan): Argo CD's applications.argoproj.io answers to a plain kubectl get applications.

Logs

The Go services (discovery, exporter, auth, notifier) write one plain-text line per event to stdout:

RollbackController: 2026/09/27 10:15:02 [ERROR] <message>

The line starts with the component, then the date and time, then the level: [DEBUG], [INFO], [WARNING] or [ERROR]. The analyzer writes <timestamp> | <level> | <module>:<function>:<line> | <message> to stderr.

kubectl logs -n telark deploy/telark-discovery-service | grep '\[ERROR\]'

Log lines carry no request ID, and there is no tracing.

What to alert on

Telark ships no alert rules. Use the Kubernetes signals your monitoring stack already collects:

  • Pod restarts, OOMKilled and failing readiness probes in the telark namespace.
  • Fill level of the telark-exporter-snapshots-pvc and telark-exporter-reports-pvc volumes.
  • Kyverno admission pods not ready. With the default fail-open setting, enforce plans stop blocking while Kyverno is down.
  • Redis memory near its limit (512 MiB by default). Redis never evicts, so it is OOM-killed when full.

The ServiceMonitor option

The chart has monitoring.serviceMonitor.enabled (off by default), which renders a Prometheus Operator ServiceMonitor for /metrics on every service. Leave it off for now: no service serves /metrics, so it collects nothing. If you later scrape from outside the namespace, the default NetworkPolicies also need an extra rule that admits your Prometheus namespace.

Next