Change history and rollback
Every change to an application is recorded field by field, deletions included, with a snapshot you can roll back to from the dashboard.
When an app breaks, the first question is "what changed?". Answering it usually means digging through events, rollout history and pod status across several workloads, and a deleted object leaves no trace to restore from.
What you get
- A change history per application, newest first: what changed, down to
the field (
Deployment/api spec.template.spec.containers[web].livenessProbe.periodSeconds: 10 → 30), when, and by whom when Telark can tell. - Each change labelled with a class (
deployment,scaling,config,incident, …) and a severity, so you can triage the feed instead of reading all of it. - A snapshot of the manifests as they were just before each change, deleted objects included.
- Rollback to any retained snapshot from the application's page, with a status you can follow and an abort while it is still pending.
How it works
Changes
Discovery compares each application's live state with the last stored
one on every informer flush. A non-empty difference becomes one entry in
status.history.changeLog with a monotonic generation, the class and
severity, the field-level changes, and the snapshot taken before it.
Deleting a tracked resource is a change like any other; Secret values
show as <redacted>.
The author comes from the telark.io/last-modified-by, -at and
-operation annotations. A Kyverno mutate policy installed by the chart
stamps them on the kinds discovery snapshots (Secrets excluded). A
/scale write, a health-only change, or a change discovery noticed on a
periodic pass carries the detection time and no author.
Change classes
A change that touches several areas collapses to one class, picked in this order:
| Priority | Class | Recorded when | Severity |
|---|---|---|---|
| 1 | drift | Three or more of topology, deployment, scaling, resources and config changed together, the shape of an out-of-band bulk edit. | high |
| 2 | topology | Resources were added to or removed from the application. | high |
| 3 | incident | Health went to degraded or down. | high; critical when down |
| 4 | recovery | Health came back to healthy after an incident. | medium |
| 5 | deployment | An image or another pod-template field changed. | medium |
| 6 | scaling | A replica count changed. | medium |
| 7 | resources | CPU or memory requests or limits changed. | medium |
| 8 | config | Ports, environment keys, ConfigMap or Secret references, Service or Ingress fields, or another field outside the pod template changed. | low |
| – | rollback | A snapshot was restored. Recorded by the rollback controller. | low |
The first time an application is recorded there is no change entry; its
baseline snapshot carries the class initial. The per-field detail stays
on the entry whatever class it gets.
Snapshots
A snapshot holds the manifests of the application's tracked resources,
one record per namespace and generation. Files are written to the
exporter's snapshot volume; the entry (id, generation,
changeClass, severity, takenAt, namespace, path) goes on the
application's status.snapshots.
Retention keeps the last 5 generations per application by default,
every namespace of a generation together. Change it under Settings
(spec.snapshots.maxPerApp on TelarkConfig); the new limit applies
from the next recorded change. Older generations are deleted from the
volume and from the status.
Secret values are stored as-is on the volume, because a rollback has to
restore them. The API masks them ([redacted]) for every signed-in user,
in the manifest view and in the compare view. Only a caller holding the
service token, which the rollback controller uses, reads them.
Rollback
- On the application's page, pick a snapshot and confirm. The dashboard holds the request for a 5-second undo window, then sends it.
- Discovery appends a
pendingentry tostatus.rollbacks, attributed to the signed-in caller. A second rollback of the same application is refused (409) while one is pending or in progress. - The rollback controller on the discovery leader picks the entry up,
marks it
in_progressand loads every namespace of the target generation. - It orders the manifests so dependencies land first:
ServiceAccount,Secret,ConfigMap,Service,NetworkPolicy, thenDeployment,StatefulSet,DaemonSet,CronJob, then any other kind.Jobmanifests are skipped, since re-applying one would start a new run. - It validates the set: every namespace exists, every API version is served, and a server-side dry run of the whole set passes.
- It replaces each object with its snapshot (created when missing, otherwise updated over the live version, keeping live owner references), so fields added after the snapshot are removed. The pass is retried up to three times, 60 seconds per attempt, within a 3-minute budget.
- It records a
rollbackchange, sets the entry tosuccessorfailed, and sends the requester an in-app notification.
The state it replaces is snapshotted first, so a rollback can itself be rolled back. A rollback to a generation that does not cover one of the application's current namespaces is refused with that namespace named.
| Status | Meaning |
|---|---|
pending | Recorded; the controller has not picked it up. Abort is still possible. |
in_progress | Loading or applying manifests. |
success | Every object was applied. |
failed | An error stopped the apply. Partial state is left as is for a person to inspect; nothing rolls forward. An entry stuck in progress for 5 minutes, after a crash for example, is marked failed. |
aborted | A user aborted it while it was pending. Nothing was applied. |
Four workers process rollbacks (DISCOVERY_ROLLBACK_WORKERS), one per
application at a time. Objects are written under the field manager
telark-discovery-service, so re-applying the same snapshot converges on
the same state.
Where it shows in the dashboard
On an application's page:
- History changes: the change log, ten entries per page, with class, severity, field-level changes and whether a snapshot backs the entry.
- Manage snapshots: every snapshot with size, time and severity, a manifest viewer, a compare mode, and rollback.
- Manage rollbacks: recorded rollbacks with status, the error behind a failed one, and Abort rollback while pending.
- Storage: usage of the snapshot volume, per application and in total. Settings shows the aggregate too.
Limits
- Snapshots cover manifests only. Volume contents, Secrets resolved from an external KMS, and objects outside the tracked kinds are out of scope. Use a backup tool such as Velero for data.
- History is bounded by retention. Keep your own copy of manifests you need beyond it.
- A failed rollback does not undo what it already applied.
- The author of a change is known only when the change went through
admission with the annotation policy in place. That policy renders
only once Kyverno's
ClusterPolicyAPI exists, so a first install gets it on its firsthelm upgrade.
Reference
Protection plans
Lock chosen changes on an application or namespace for a set time window, prove the lock is in force, and keep a report of what it blocked.
Insights
Incident decision support. One card per affected workload with the likely cause, the evidence and the change it followed, plus a setup review of 60 rules. Local, read-only and optional.