Guides

Run discovery with several replicas

Scale the discovery service for load and failover, and check which replica is the leader.

Discovery can run several replicas. They share the per-application work and the API traffic, while one elected leader runs the work that must happen once. If the leader dies, another replica takes over within about 15 seconds.

What runs where

WorkRuns on
Per-application jobs: derive, diff, snapshot and publish each applicationEvery replica, one application at a time under a Redis lock
Informer cachesEvery replica, so any of them can become leader
HTTP APIEvery replica
Recording informer changes, the periodic discovery pass, force sync, the protection plan controller, report checkpoints, the rollback controllerThe leader only

Every replica holds the full informer cache, so memory per replica grows with cluster size.

1. Choose the replica count

In standard and performance, an autoscaler already scales discovery on CPU from 1 up to 3 or 5 replicas. minimal runs one replica with no failover.

To keep at least two replicas at all times, set a floor:

helm upgrade telark oci://ghcr.io/telark/charts/telark -n telark <your install flags> \
  --set services.discovery.autoscaling.minReplicas=2

In performance, a disruption budget also keeps one pod running through node drains.

2. Check the replicas and the leader

kubectl get pods -n telark -l app.kubernetes.io/component=discovery-service
kubectl exec -n telark telark-redis-master-0 -- redis-cli GET election:prewarm

You should see the expected number of Running pods, and the second command should print the name of one of them. That pod is the leader. The key is a 15-second lease the leader renews every 5 seconds (COORDINATION_ELECTION_TTL_SEC, COORDINATION_ELECTION_RENEW_SEC).

3. Test failover

kubectl delete pod -n telark <leader-pod>

Within about 15 seconds, GET election:prewarm should name another pod. The other replicas keep serving the API throughout, and plan transitions and rollbacks resume on the new leader.

Signals to watch

Telark exports no Prometheus metrics yet, so watch these:

  • The leader lease. election:prewarm should never stay empty longer than one lease.
  • The discovery pass. GET /api/v1/discovery/status on discovery returns remaining, the application jobs not yet processed. A backlog that keeps growing between passes means the replicas are saturated.
  • Force sync. status.lastForceSync.phase on the Application (queued, running, completed or failed).
  • Logs. Redis or exporter errors appear as [ERROR] lines in the discovery logs.

Tuning the plan controller

The protection plan controller runs on the leader. Plans start and end on a timer set to their window edges. Separately, PROTECTION_PLAN_TICK_INTERVAL_SEC (31 seconds by default) sets how often the leader checks plan health and repairs policies. Lowering it catches drift sooner and adds load on the leader:

--set-string services.discovery.env.PROTECTION_PLAN_TICK_INTERVAL_SEC=15

Limits

  • Replicas share work within one cluster. They do not make Telark multi-cluster.
  • The exporter's snapshot volume is not replicated by discovery replicas. See Back up Telark.