Get started
Architecture overview
The services one Helm release installs, and how a change moves through them.
Everything runs inside your cluster from one Helm release. Five Go and Python services do the work; the dashboard is a static web app in front of them.
Services
| Service | What it does |
|---|---|
discovery | Watches workloads and groups them into applications, records changes and snapshots, runs rollbacks, and drives the protection plan lifecycle: approvals, Kyverno policies, health checks, violations and reports. Leader-elected. |
exporter | The only writer of Telark custom resources. Stores snapshots and plan reports on its two volumes and seeds the built-in roles, categories and the default TelarkConfig. |
auth | Passkey and Google sign-in, sessions, and cleanup when a user, group or role is deleted. |
notifier | Consumes application change events from NATS and persists them through the exporter. |
analyzer | Insights: incident cards and setup recommendations, using read-only cluster access and a local model served by Ollama. Optional. |
ui | The dashboard. |
The chart also installs Kyverno (admission enforcement), Redis (leader locks, queues, caches), NATS JetStream (the event stream), metrics-server and Ollama.
How a change flows
- Discover. Discovery sees a workload change, groups it into its application, compares it with the stored state, and classifies it (deployment, scaling, config, incident and so on). It stores the pre-change manifests as a snapshot through the exporter and publishes the change on NATS; the notifier writes the
Applicationresource. - Protect. When a protection plan's window opens, discovery renders its templates into namespaced Kyverno
Policyobjects labelledtelark.io/protection-plan=<plan-id>, in audit or enforce mode. - Verify. On every controller tick (31 seconds by default), discovery reads those policies back from the cluster, marks the plan healthy, drifted or degraded, and redeploys anything that drifted. Violations come from the Kubernetes events Kyverno raises.
- Explain. An incident or recovery queues an analysis job in Redis. When Insights is on, the analyzer runs its rules, has the local model rewrite the wording, and the dashboard shows the result.
When a plan ends, discovery deletes its policies and stores a final report through the exporter.
Where state lives
| Store | Holds |
|---|---|
Custom resources (telark.io/v1alpha1, in etcd) | Applications, protection plans, settings, users, groups, roles, passkeys, sessions. The system of record. |
| Snapshot volume | Pre-change manifests, one file per application, namespace and generation. |
| Reports volume | Protection plan reports and their violation ledgers. |
| Redis | Coordination, queues, caches, in-app notifications and insights. Treat it as disposable. |
Security boundaries
- Only the exporter writes Telark custom resources; discovery may also patch
applications/statusfor rollbacks. The CRD write guard, on by default, rejects writes from any other identity at admission. - Users authenticate with a session token; services with a shared service token generated by the chart.
- Ingress NetworkPolicies, on by default, let only Telark pods of the same release reach the service APIs, and only discovery and the notifier reach NATS.
- The analyzer's Kubernetes access is read-only. Auth, the notifier and the dashboard have none.