Run discovery with several replicas
Scale the discovery service for load and failover, and check which replica is the leader.
Discovery can run several replicas. They share the per-application work and the API traffic, while one elected leader runs the work that must happen once. If the leader dies, another replica takes over within about 15 seconds.
What runs where
| Work | Runs on |
|---|---|
| Per-application jobs: derive, diff, snapshot and publish each application | Every replica, one application at a time under a Redis lock |
| Informer caches | Every replica, so any of them can become leader |
| HTTP API | Every replica |
| Recording informer changes, the periodic discovery pass, force sync, the protection plan controller, report checkpoints, the rollback controller | The leader only |
Every replica holds the full informer cache, so memory per replica grows with cluster size.
1. Choose the replica count
In standard and performance, an autoscaler already scales discovery on CPU from 1 up to 3 or 5 replicas. minimal runs one replica with no failover.
To keep at least two replicas at all times, set a floor:
helm upgrade telark oci://ghcr.io/telark/charts/telark -n telark <your install flags> \
--set services.discovery.autoscaling.minReplicas=2In performance, a disruption budget also keeps one pod running through node drains.
2. Check the replicas and the leader
kubectl get pods -n telark -l app.kubernetes.io/component=discovery-service
kubectl exec -n telark telark-redis-master-0 -- redis-cli GET election:prewarmYou should see the expected number of Running pods, and the second command should print the name of one of them. That pod is the leader. The key is a 15-second lease the leader renews every 5 seconds (COORDINATION_ELECTION_TTL_SEC, COORDINATION_ELECTION_RENEW_SEC).
3. Test failover
kubectl delete pod -n telark <leader-pod>Within about 15 seconds, GET election:prewarm should name another pod. The other replicas keep serving the API throughout, and plan transitions and rollbacks resume on the new leader.
Signals to watch
Telark exports no Prometheus metrics yet, so watch these:
- The leader lease.
election:prewarmshould never stay empty longer than one lease. - The discovery pass.
GET /api/v1/discovery/statuson discovery returnsremaining, the application jobs not yet processed. A backlog that keeps growing between passes means the replicas are saturated. - Force sync.
status.lastForceSync.phaseon theApplication(queued,running,completedorfailed). - Logs. Redis or exporter errors appear as
[ERROR]lines in the discovery logs.
Tuning the plan controller
The protection plan controller runs on the leader. Plans start and end on a timer set to their window edges. Separately, PROTECTION_PLAN_TICK_INTERVAL_SEC (31 seconds by default) sets how often the leader checks plan health and repairs policies. Lowering it catches drift sooner and adds load on the leader:
--set-string services.discovery.env.PROTECTION_PLAN_TICK_INTERVAL_SEC=15Limits
- Replicas share work within one cluster. They do not make Telark multi-cluster.
- The exporter's snapshot volume is not replicated by discovery replicas. See Back up Telark.