Features

Change history and rollback

Every change to an application is recorded field by field, deletions included, with a snapshot you can roll back to from the dashboard.

When an app breaks, the first question is "what changed?". Answering it usually means digging through events, rollout history and pod status across several workloads, and a deleted object leaves no trace to restore from.

What you get

  • A change history per application, newest first: what changed, down to the field (Deployment/api spec.template.spec.containers[web].livenessProbe.periodSeconds: 10 → 30), when, and by whom when Telark can tell.
  • Each change labelled with a class (deployment, scaling, config, incident, …) and a severity, so you can triage the feed instead of reading all of it.
  • A snapshot of the manifests as they were just before each change, deleted objects included.
  • Rollback to any retained snapshot from the application's page, with a status you can follow and an abort while it is still pending.

How it works

Changes

Discovery compares each application's live state with the last stored one on every informer flush. A non-empty difference becomes one entry in status.history.changeLog with a monotonic generation, the class and severity, the field-level changes, and the snapshot taken before it. Deleting a tracked resource is a change like any other; Secret values show as <redacted>.

The author comes from the telark.io/last-modified-by, -at and -operation annotations. A Kyverno mutate policy installed by the chart stamps them on the kinds discovery snapshots (Secrets excluded). A /scale write, a health-only change, or a change discovery noticed on a periodic pass carries the detection time and no author.

Change classes

A change that touches several areas collapses to one class, picked in this order:

PriorityClassRecorded whenSeverity
1driftThree or more of topology, deployment, scaling, resources and config changed together, the shape of an out-of-band bulk edit.high
2topologyResources were added to or removed from the application.high
3incidentHealth went to degraded or down.high; critical when down
4recoveryHealth came back to healthy after an incident.medium
5deploymentAn image or another pod-template field changed.medium
6scalingA replica count changed.medium
7resourcesCPU or memory requests or limits changed.medium
8configPorts, environment keys, ConfigMap or Secret references, Service or Ingress fields, or another field outside the pod template changed.low
–rollbackA snapshot was restored. Recorded by the rollback controller.low

The first time an application is recorded there is no change entry; its baseline snapshot carries the class initial. The per-field detail stays on the entry whatever class it gets.

Snapshots

A snapshot holds the manifests of the application's tracked resources, one record per namespace and generation. Files are written to the exporter's snapshot volume; the entry (id, generation, changeClass, severity, takenAt, namespace, path) goes on the application's status.snapshots.

Retention keeps the last 5 generations per application by default, every namespace of a generation together. Change it under Settings (spec.snapshots.maxPerApp on TelarkConfig); the new limit applies from the next recorded change. Older generations are deleted from the volume and from the status.

Secret values are stored as-is on the volume, because a rollback has to restore them. The API masks them ([redacted]) for every signed-in user, in the manifest view and in the compare view. Only a caller holding the service token, which the rollback controller uses, reads them.

Rollback

  1. On the application's page, pick a snapshot and confirm. The dashboard holds the request for a 5-second undo window, then sends it.
  2. Discovery appends a pending entry to status.rollbacks, attributed to the signed-in caller. A second rollback of the same application is refused (409) while one is pending or in progress.
  3. The rollback controller on the discovery leader picks the entry up, marks it in_progress and loads every namespace of the target generation.
  4. It orders the manifests so dependencies land first: ServiceAccount, Secret, ConfigMap, Service, NetworkPolicy, then Deployment, StatefulSet, DaemonSet, CronJob, then any other kind. Job manifests are skipped, since re-applying one would start a new run.
  5. It validates the set: every namespace exists, every API version is served, and a server-side dry run of the whole set passes.
  6. It replaces each object with its snapshot (created when missing, otherwise updated over the live version, keeping live owner references), so fields added after the snapshot are removed. The pass is retried up to three times, 60 seconds per attempt, within a 3-minute budget.
  7. It records a rollback change, sets the entry to success or failed, and sends the requester an in-app notification.

The state it replaces is snapshotted first, so a rollback can itself be rolled back. A rollback to a generation that does not cover one of the application's current namespaces is refused with that namespace named.

StatusMeaning
pendingRecorded; the controller has not picked it up. Abort is still possible.
in_progressLoading or applying manifests.
successEvery object was applied.
failedAn error stopped the apply. Partial state is left as is for a person to inspect; nothing rolls forward. An entry stuck in progress for 5 minutes, after a crash for example, is marked failed.
abortedA user aborted it while it was pending. Nothing was applied.

Four workers process rollbacks (DISCOVERY_ROLLBACK_WORKERS), one per application at a time. Objects are written under the field manager telark-discovery-service, so re-applying the same snapshot converges on the same state.

Where it shows in the dashboard

On an application's page:

  • History changes: the change log, ten entries per page, with class, severity, field-level changes and whether a snapshot backs the entry.
  • Manage snapshots: every snapshot with size, time and severity, a manifest viewer, a compare mode, and rollback.
  • Manage rollbacks: recorded rollbacks with status, the error behind a failed one, and Abort rollback while pending.
  • Storage: usage of the snapshot volume, per application and in total. Settings shows the aggregate too.

Limits

  • Snapshots cover manifests only. Volume contents, Secrets resolved from an external KMS, and objects outside the tracked kinds are out of scope. Use a backup tool such as Velero for data.
  • History is bounded by retention. Keep your own copy of manifests you need beyond it.
  • A failed rollback does not undo what it already applied.
  • The author of a change is known only when the change went through admission with the annotation policy in place. That policy renders only once Kyverno's ClusterPolicy API exists, so a first install gets it on its first helm upgrade.

Reference