Features

Protection plans

Lock chosen changes on an application or namespace for a set time window, prove the lock is in force, and keep a report of what it blocked.

In a shared cluster, anyone with write access can delete, scale or re-image a critical app at any time. RBAC is permanent and per object; it cannot say "nobody changes checkout's images or Secrets between 22:00 and 02:00 tonight". A change freeze announced in chat depends on everyone reading it.

What you get

  • A plan that locks specific kinds of change (creation, deletion, image changes, scaling, ConfigMap and Secret edits, storage, …) for chosen applications or namespaces.
  • A time window: the plan arms itself at its start and disarms itself at its end, or stays on until you cancel it.
  • Audit mode to see what would have been blocked, then enforce mode to block it.
  • Optional approval before anything is enforced. Plans in the Production environment always require one.
  • Health checked against the live cluster, with drift repaired while the plan is active.
  • A violations feed while the plan runs, and a downloadable report (HTML, Markdown, JSON or CSV) when it ends.

Plans live on the Protection plans page of the dashboard.

How it works

A plan (ProtectionPlan, protectionplans.telark.io) has three parts:

  1. Scope. A list of applications, or a list of namespaces. Exclusions carve kinds, or named resources of the selected applications, back out.
  2. Templates. One or more policy templates from the built-in catalog, with parameters where a template needs them.
  3. Window. permanent (until canceled) or time_range with startAt and endAt.

While the plan is active, discovery renders each template into namespaced Kyverno Policy objects, one per template and namespace, and applies them. Kyverno ships with the chart and enforces them at admission. When the plan ends or is canceled, those policies are deleted. Nothing stays in the cluster outside the window.

Templates

TemplateBlocksScope
block-createCreating any resourceapplications, namespaces
block-updateUpdating any resource, including /scaleapplications, namespaces
block-deleteDeleting any resourceapplications, namespaces
block-image-typesImages matching glob patterns (imagePatterns)applications, namespaces
block-image-tagsImages with listed tags (tags)applications, namespaces
block-replica-scalingReplica changes on Deployments and StatefulSetsapplications only
block-storage-changesPVC create, update, delete, and workload volume changesapplications, namespaces
block-config-secret-resource-changesUpdating or deleting ConfigMaps and Secretsapplications, namespaces
block-workload-config-mount-changesChanges to volume mounts and ConfigMap or Secret sources in workloadsapplications, namespaces

Rule details are in Policy templates. The catalog is fixed; you cannot author your own templates today.

Audit and enforce

The plan's mode sets validationFailureAction on every policy it deploys.

  • audit: the change is admitted and recorded as a violation. The message reads "would be blocked by protection plan …".
  • enforce: the change is refused. The caller sees the Kubernetes API error at the source of the change.

Enforcing a namespace scope needs Owner on protection-plans; a Contributor can create audit plans on namespaces and enforce plans on applications.

Lifecycle

PhaseMeaningPolicies in the cluster
pending_approvalWaiting for an approver.None
scheduledstartAt is in the future.None
activeIn force. Health and violations are live.Deployed
terminatedThe window ended, either while active or while still waiting for approval.Removed
canceledCanceled by a user, or rejected by an approver.Removed
failedActivation failed (scope resolution, render or deploy).Removed

A controller on the discovery leader moves plans across window edges. It runs every 31 seconds by default (PROTECTION_PLAN_TICK_INTERVAL_SEC) plus a timer set on the next window edge. You can:

  • cancel an active, scheduled, pending_approval or failed plan;
  • reactivate a terminated, canceled or failed plan, which re-enters the lifecycle (and asks for approval again where required);
  • edit a plan in any phase. Changing the window recomputes the phase; an active plan given a future start withdraws its policies until then. An edit of a canceled or terminated plan takes effect on reactivation. The scope type cannot change.
  • duplicate a plan into a new one.

Exclusions

  • Kinds: exclude whole kinds, for example ConfigMap, from either scope type.
  • Resources: exclude named resources (kind, name, namespace) of the selected applications. Applications scope only.

Every rendered rule carries the exclusions, whichever template produced it; excluding a kind or resource also excludes its /scale subresource. Changing exclusions re-renders the plan's policies and follows the same approval rules as any scope change.

Approvals

The server derives approvalMode when a plan is created or duplicated:

  • A plan in the Production environment is always required.
  • Elsewhere the default is automatic. A client may ask for required; asking for automatic counts only from an Owner on protection-plans.
  • The mode cannot change afterwards, and an edit may not move an automatic plan into Production.

A required plan waits in pending_approval with nothing deployed. Every user who may approve it gets an in-app notification. An approver (Owner on protection-plans) either approves it, and the plan becomes scheduled or active, or rejects it with a mandatory comment, and the plan is canceled.

  • Nobody who put the current version up for approval can decide it: the creator, and anyone who reactivated or materially edited it since the last approval.
  • A decision names the request it answers, so a stale or concurrent decision is refused (409).
  • Once approved, material edits (scope, exclusions, templates, mode, window) are refused. Duplicate the plan, or cancel and reactivate it, to request a new approval. A material edit of a plan still pending asks for approval again.

Environments and tags

A plan carries at most one environment and up to 20 tags, both categories. Production, Staging and Development ship as environments; Compliance, Security and Baseline as tags. Add your own under Organize on the Protection plans page. They filter the plan list and never change what a plan renders, except that Production requires approval.

Health

Health is read from the cluster, not assumed from the API write. On each controller tick, discovery lists the Kyverno policies labelled telark.io/protection-plan=<plan id> and compares each with a fresh render of the plan (the telark.io/render-hash annotation).

HealthMeaning
healthyEvery declared policy is present, ready, current, and uses the plan's mode.
driftedA declared policy is missing or edited, its failure action no longer matches the mode, or an unexpected policy carries the plan label.
degradedA policy is present but Kyverno has not made it ready. Wins over drift.
unknownThe plan is not active.

While the plan is active, the same tick repairs drift: missing or outdated policies are redeployed, a changed failure action is patched back, and unexpected policies with the plan label are deleted. Managed policies left behind by a plan that ended are swept too. Refresh health in the dashboard computes health on request without repairing or storing it. The result is stored on the plan's status, with the conditions Ready, Approved and PoliciesHealthy.

Violations

Violations are the Kyverno PolicyViolation events of the plan's policies, across the namespaces its scope resolves to, newest first. Each carries the policy, rule, resource (kind, name, namespace), result (pass, fail, warn, error, skip), message and timestamp. The feed returns 50 by default and 200 at most.

Kubernetes keeps events for one hour by default, so the live feed shows recent decisions only. A finished plan shows none: its policies are gone. Reports keep the full record.

Reports

A report covers what happened while a plan was in effect: its configuration, scope and policies, timeline, health over the run, and the violations it recorded.

  • A final report is captured automatically when an active plan ends or is canceled. A plan that never started has none.
  • Generate one on demand from the plan's page once the plan has started. Two renders run at a time; a third answers 429 with Retry-After. The newest 10 on-demand reports per plan are kept; final reports are never pruned.
  • While a plan runs, its violations are checkpointed into a per-plan ledger every 15 minutes (PROTECTION_PLAN_REPORT_CHECKPOINT_SEC), so a report is not limited to the one-hour event window. A report keeps at most 5,000 violations and says when it was truncated.
  • The Reports tab lists every plan's reports, filterable by plan, environment, trigger (manual, cancel, end) and date.
  • Reports live on the exporter's reports volume and are deleted with their plan.

Permissions

Every plan action has its own permission on the protection-plans scope, and a custom role can withhold any one of them with protection-plans.<action>.deny.

ActionMinimum level
viewprotectionplans, viewprotectionplanviolations, viewprotectionplanreports, downloadprotectionplanreportReadOnly
createprotectionplan, editprotectionplan, duplicateprotectionplan, cancelprotectionplan, reactivateprotectionplan, generateprotectionplanreport, addprotectionplancategoryContributor
approveprotectionplan, rejectprotectionplan, deleteprotectionplan, editprotectionplancategory, deleteprotectionplancategoryOwner

The three *protectionplancategory actions govern environments and tags. Report files include admission decisions, so withholding violations alone does not hide them.

Validation

Discovery refuses a plan when:

  • its name is taken, ignoring case and surrounding spaces (409);
  • its name or description contains {{ or }} (400);
  • its window has already ended (400);
  • it targets the release namespace, an excluded namespace or a namespace that does not exist (400); while the excluded list cannot be read, a namespace-scoped plan answers 503;
  • a template does not support the scope, is listed twice, or gets a parameter it does not declare (400).

Names are at most 64 characters and descriptions 512. Request bodies are capped at 1 MiB (413) and unknown fields are refused (400).

Limits

  • Kyverno fails open by default (app.kyverno.failOpen=true): while its webhook is down, changes are admitted even for enforcing plans. Set app.kyverno.failOpen=false together with kyverno.features.forceFailurePolicyIgnore.enabled=false to fail closed.
  • Plans act at admission, inside the cluster. They do not gate CI pipelines, and they cannot undo a change made before the plan started.
  • Templates are a fixed catalog.
  • Live violations expire with Kubernetes events (one hour by default). Use reports for the history of a finished plan.

Reference