Features

Insights

Incident decision support. One card per affected workload with the likely cause, the evidence and the change it followed, plus a setup review of 60 rules. Local, read-only and optional.

When an application degrades, the on-call engineer has to piece together events, pod status and recent changes across several workloads before deciding what to do. Insights does that first pass for you and shows its evidence, so you can check it.

Insights is on by default and optional. It is decision support, not autopilot. It never changes anything in the cluster, and its output can be wrong: confirm it before you act.

What you get

Incident cards. When an application degrades, Insights writes one card per affected workload with:

  • the likely cause: crash loop, out-of-memory kill, image pull failure, scheduling failure, failing probe, resource pressure, stuck rollout, or a regression after a change, with a precise sub-cause (for example image not found rather than image pull);
  • the evidence it used: events, pod and container status, exit codes, restart counts;
  • the recorded change it followed, when the problem started within 30 minutes of one, with a pointer to the snapshot taken before it;
  • up to four concrete next steps.

Setup review. Each application's configuration is checked against 60 deterministic rules in 10 families, and each finding is listed as a recommendation:

FamilyExample finding
ReliabilityA Deployment runs a single replica, or has replicas but no disruption budget.
ResourcesA container has no requests, or runs close to its memory limit.
ScalingAn autoscaler is held at its maximum, or cannot read utilization.
SecurityA container runs privileged or as root.
ImagesA container uses latest or no tag.
ConfigurationThe same environment variable is defined twice.
NetworkingA Service selects no pods, or targets a port no container declares.
Change riskTwo or more rollbacks in the last seven days.
ProtectionA production application has no protection plan, or is only audited.
ConsistencyThe same workload runs different images in two namespaces.

The full list is in Insight rules.

Both appear on the Insights page, and on each application's page.

How it works

Deterministic rules decide

The analyzer reads the application with four read-only tools: its overview, its recent change history, its warning events of the last hour, and the status of up to three workloads (those named by events first). Rules turn those facts into cards. The rules set every card's kind, cause, severity, confidence, evidence and state. A workload with no ready replica is Critical, whatever the cause. A card is ready about one to two seconds after the run starts.

A liveness or startup probe that keeps restarting a container is reported as a crash loop, since the restarts are what take the workload down. A paused Deployment is not an incident unless its pods fail.

A local model only rewrites the wording

In the default fast mode, a small open-weight model served by Ollama in your cluster then rewrites each card's title and summary in plain sentences, in one call with no tools. Telark keeps a rewrite only if it names the workload, keeps every name and number of the rule text and the phrase that names the cause, and adds no number and no symptom of another kind of problem. Otherwise the card keeps the rule-written text and the run is marked rules only. The rule text is complete on its own.

The setup review never calls the model.

Fast and deep mode

Fast (default)Deep (opt-in)
Who investigatesRulesThe model, through the same four read-only tools, over up to 8 steps
Model neededAny installed model; runs without one (rules only)A model with tool calling, loaded and ready
Time on a 2 vCPU nodeCards in 1–2 s, narration in about 8 s per insightMinutes per run
HardwareCPUA GPU in practice

Switch with --set services.analyzer.env.ANALYZER_MODE=deep.

Hardware profiles

The chart's Ollama runtime is sized once for every install mode.

ProfileModelRuntime resourcesChart valuesLatency
CPU tiny (default)granite4:350m (708 MB)request 250m CPU and 1536Mi, limit 2 CPUdefaultsabout 8 s per insight once loaded
CPU, 4 vCPUqwen3:1.7brequest 1 CPU and 4Gi, limit 4 CPUollama.resources.*, services.analyzer.env.ANALYZER_NUM_THREAD=420–45 s
GPU, deepqwen3:4brequest 1 CPU and 3Gi, one nvidia.com/gpuollama.ollama.gpu.enabled=true, ANALYZER_MODE=deep, ANALYZER_CONTEXT_TOKENS=8192, OLLAMA_CONTEXT_LENGTH=8192minutes on CPU, seconds on GPU

Keep ANALYZER_NUM_THREAD equal to the runtime's CPU limit and never above the node's vCPU count; more threads than cores makes narration many times slower. The default fits the minimal install on one 2 vCPU, 8 GiB node. The first run after a runtime restart takes a few seconds longer while the model loads.

When analysis runs

  • On demand: Analyze on an application in the Insights page or on its page.
  • Automatically: turn on analyze automatically in Settings, and discovery queues a run each time it records an incident or a recovery. A recovery resolves the open cards without calling the model.
  • Setup review: after every run, and from a background sweep that reviews changed applications first, then any application not reviewed for two hours, at most 20 per minute.

The Insights page

The Insights entry sits right after Applications in the sidebar, with two tabs, Incidents and Recommendations. You can filter by severity, state, kind or family, namespace and environment, search, sort, and group by application, namespace or category. A card opens a details panel with the cause, next steps, the facts behind it and the evidence.

  • States: Open, Updated, Resolved, and Stale (active but not seen for a day).
  • Triage: Acknowledge marks a card as seen; Dismiss hides a recommendation that does not apply until its facts change; Reopen clears either.
  • The page refreshes every 15 seconds while visible. It is served by discovery, so it stays readable while the analyzer is off.

Settings

The switch, the model and automatic analysis live on the TelarkConfig resource and are edited under Settings (Owner on settings):

spec:
  ai:
    enabled: true           # fresh-install default
    model: granite4:350m    # fresh-install default
    autoAnalyze: false      # fresh-install default

These defaults are written only when the TelarkConfig is first created; an upgrade keeps your settings.

Where the model comes from

  • Connected (app.ollama.autoPull=true, default): the analyzer asks the in-cluster runtime to download the chosen model when it lacks it. Only models whose licence Telark knows are pulled; every model the dashboard lists is Apache-2.0. Only the runtime gets outbound HTTPS.
  • Air-gapped (app.ollama.autoPull=false): nothing is downloaded and the runtime has no egress. Load the model onto the runtime's volume yourself, then pick it in Settings.
  • Your own runtime (app.ollama.runtimeUrl): point the analyzer at an Ollama-API server you run, for example on a GPU host, and set app.ollama.enabled=false. No key. The facts of each run are sent to that endpoint, so keep it on infrastructure you control.
helm upgrade telark oci://ghcr.io/telark/charts/telark -n telark \
  --set app.ollama.runtimeUrl=http://ollama.gpu.internal:11434 \
  --set app.ollama.enabled=false

Safety boundaries

  • Read-only. The analyzer has no write tool. Its ClusterRole grants only get and list on pods, events, Deployments, StatefulSets, DaemonSets, ReplicaSets, Services, PodDisruptionBudgets, HorizontalPodAutoscalers and NetworkPolicies. It never reads Secret or ConfigMap contents, nodes, RBAC objects or logs; rules that look at environment variables store their names, never their values.
  • No data leaves the cluster. Models are open-weight and run on Ollama, an open-source runtime. There is no cloud provider and no API key, in any mode. A network policy admits only the analyzer to the runtime.
  • Code owns the facts. Ids, kinds, severities, evidence, timestamps and whether a card is open or resolved are set by Telark code, never by the model.
  • Respects your scope. Excluded namespaces are never analyzed or shown, and every request checks the caller's permissions on the insights scope.
  • On by default, and optional. A fresh install enables it with granite4:350m; setup reviews run on their own, and incident analysis runs when you click Analyze or once automatic analysis is on. Turn it off in Settings, or skip the runtime with app.ollama.enabled=false. Protection plans, history and rollback never wait on it; if the analyzer is down, discovery keeps serving the last cards.

Limits

  • Output can be wrong. The cause is a best guess from the evidence shown, and a rule that relies on a heuristic says so with Medium confidence.
  • At most three incident cards per run, and up to 10 workloads per setup review. An application keeps its 40 most severe recommendations.
  • The usage rules (near a limit, over- or under-provisioned) need metrics-server and at least 12 samples over 12 hours.
  • Deep mode needs a GPU to be practical, and records a failed run when the model is missing or times out.
  • Some checks are not available yet, among them zone topology spread, measured CPU throttling, image vulnerability scanning, and ConfigMap changes without a rollout.

Reference