Insights
Incident decision support. One card per affected workload with the likely cause, the evidence and the change it followed, plus a setup review of 60 rules. Local, read-only and optional.
When an application degrades, the on-call engineer has to piece together events, pod status and recent changes across several workloads before deciding what to do. Insights does that first pass for you and shows its evidence, so you can check it.
Insights is on by default and optional. It is decision support, not autopilot. It never changes anything in the cluster, and its output can be wrong: confirm it before you act.
What you get
Incident cards. When an application degrades, Insights writes one card per affected workload with:
- the likely cause: crash loop, out-of-memory kill, image pull failure, scheduling failure, failing probe, resource pressure, stuck rollout, or a regression after a change, with a precise sub-cause (for example image not found rather than image pull);
- the evidence it used: events, pod and container status, exit codes, restart counts;
- the recorded change it followed, when the problem started within 30 minutes of one, with a pointer to the snapshot taken before it;
- up to four concrete next steps.
Setup review. Each application's configuration is checked against 60 deterministic rules in 10 families, and each finding is listed as a recommendation:
| Family | Example finding |
|---|---|
| Reliability | A Deployment runs a single replica, or has replicas but no disruption budget. |
| Resources | A container has no requests, or runs close to its memory limit. |
| Scaling | An autoscaler is held at its maximum, or cannot read utilization. |
| Security | A container runs privileged or as root. |
| Images | A container uses latest or no tag. |
| Configuration | The same environment variable is defined twice. |
| Networking | A Service selects no pods, or targets a port no container declares. |
| Change risk | Two or more rollbacks in the last seven days. |
| Protection | A production application has no protection plan, or is only audited. |
| Consistency | The same workload runs different images in two namespaces. |
The full list is in Insight rules.
Both appear on the Insights page, and on each application's page.
How it works
Deterministic rules decide
The analyzer reads the application with four read-only tools: its overview, its recent change history, its warning events of the last hour, and the status of up to three workloads (those named by events first). Rules turn those facts into cards. The rules set every card's kind, cause, severity, confidence, evidence and state. A workload with no ready replica is Critical, whatever the cause. A card is ready about one to two seconds after the run starts.
A liveness or startup probe that keeps restarting a container is reported as a crash loop, since the restarts are what take the workload down. A paused Deployment is not an incident unless its pods fail.
A local model only rewrites the wording
In the default fast mode, a small open-weight model served by Ollama in your cluster then rewrites each card's title and summary in plain sentences, in one call with no tools. Telark keeps a rewrite only if it names the workload, keeps every name and number of the rule text and the phrase that names the cause, and adds no number and no symptom of another kind of problem. Otherwise the card keeps the rule-written text and the run is marked rules only. The rule text is complete on its own.
The setup review never calls the model.
Fast and deep mode
| Fast (default) | Deep (opt-in) | |
|---|---|---|
| Who investigates | Rules | The model, through the same four read-only tools, over up to 8 steps |
| Model needed | Any installed model; runs without one (rules only) | A model with tool calling, loaded and ready |
| Time on a 2 vCPU node | Cards in 1–2 s, narration in about 8 s per insight | Minutes per run |
| Hardware | CPU | A GPU in practice |
Switch with --set services.analyzer.env.ANALYZER_MODE=deep.
Hardware profiles
The chart's Ollama runtime is sized once for every install mode.
| Profile | Model | Runtime resources | Chart values | Latency |
|---|---|---|---|---|
| CPU tiny (default) | granite4:350m (708 MB) | request 250m CPU and 1536Mi, limit 2 CPU | defaults | about 8 s per insight once loaded |
| CPU, 4 vCPU | qwen3:1.7b | request 1 CPU and 4Gi, limit 4 CPU | ollama.resources.*, services.analyzer.env.ANALYZER_NUM_THREAD=4 | 20–45 s |
| GPU, deep | qwen3:4b | request 1 CPU and 3Gi, one nvidia.com/gpu | ollama.ollama.gpu.enabled=true, ANALYZER_MODE=deep, ANALYZER_CONTEXT_TOKENS=8192, OLLAMA_CONTEXT_LENGTH=8192 | minutes on CPU, seconds on GPU |
Keep ANALYZER_NUM_THREAD equal to the runtime's CPU limit and never
above the node's vCPU count; more threads than cores makes narration many
times slower. The default fits the minimal install on one 2 vCPU,
8 GiB node. The first run after a runtime restart takes a few seconds
longer while the model loads.
When analysis runs
- On demand: Analyze on an application in the Insights page or on its page.
- Automatically: turn on analyze automatically in Settings, and discovery queues a run each time it records an incident or a recovery. A recovery resolves the open cards without calling the model.
- Setup review: after every run, and from a background sweep that reviews changed applications first, then any application not reviewed for two hours, at most 20 per minute.
The Insights page
The Insights entry sits right after Applications in the sidebar, with two tabs, Incidents and Recommendations. You can filter by severity, state, kind or family, namespace and environment, search, sort, and group by application, namespace or category. A card opens a details panel with the cause, next steps, the facts behind it and the evidence.
- States: Open, Updated, Resolved, and Stale (active but not seen for a day).
- Triage: Acknowledge marks a card as seen; Dismiss hides a recommendation that does not apply until its facts change; Reopen clears either.
- The page refreshes every 15 seconds while visible. It is served by discovery, so it stays readable while the analyzer is off.
Settings
The switch, the model and automatic analysis live on the TelarkConfig
resource and are edited under Settings (Owner on settings):
spec:
ai:
enabled: true # fresh-install default
model: granite4:350m # fresh-install default
autoAnalyze: false # fresh-install defaultThese defaults are written only when the TelarkConfig is first
created; an upgrade keeps your settings.
Where the model comes from
- Connected (
app.ollama.autoPull=true, default): the analyzer asks the in-cluster runtime to download the chosen model when it lacks it. Only models whose licence Telark knows are pulled; every model the dashboard lists is Apache-2.0. Only the runtime gets outbound HTTPS. - Air-gapped (
app.ollama.autoPull=false): nothing is downloaded and the runtime has no egress. Load the model onto the runtime's volume yourself, then pick it in Settings. - Your own runtime (
app.ollama.runtimeUrl): point the analyzer at an Ollama-API server you run, for example on a GPU host, and setapp.ollama.enabled=false. No key. The facts of each run are sent to that endpoint, so keep it on infrastructure you control.
helm upgrade telark oci://ghcr.io/telark/charts/telark -n telark \
--set app.ollama.runtimeUrl=http://ollama.gpu.internal:11434 \
--set app.ollama.enabled=falseSafety boundaries
- Read-only. The analyzer has no write tool. Its ClusterRole grants
only
getandliston pods, events, Deployments, StatefulSets, DaemonSets, ReplicaSets, Services, PodDisruptionBudgets, HorizontalPodAutoscalers and NetworkPolicies. It never reads Secret or ConfigMap contents, nodes, RBAC objects or logs; rules that look at environment variables store their names, never their values. - No data leaves the cluster. Models are open-weight and run on Ollama, an open-source runtime. There is no cloud provider and no API key, in any mode. A network policy admits only the analyzer to the runtime.
- Code owns the facts. Ids, kinds, severities, evidence, timestamps and whether a card is open or resolved are set by Telark code, never by the model.
- Respects your scope. Excluded namespaces are never analyzed or
shown, and every request checks the caller's permissions on the
insightsscope. - On by default, and optional. A fresh install enables it with
granite4:350m; setup reviews run on their own, and incident analysis runs when you click Analyze or once automatic analysis is on. Turn it off in Settings, or skip the runtime withapp.ollama.enabled=false. Protection plans, history and rollback never wait on it; if the analyzer is down, discovery keeps serving the last cards.
Limits
- Output can be wrong. The cause is a best guess from the evidence shown, and a rule that relies on a heuristic says so with Medium confidence.
- At most three incident cards per run, and up to 10 workloads per setup review. An application keeps its 40 most severe recommendations.
- The usage rules (near a limit, over- or under-provisioned) need metrics-server and at least 12 samples over 12 hours.
- Deep mode needs a GPU to be practical, and records a failed run when the model is missing or times out.
- Some checks are not available yet, among them zone topology spread, measured CPU throttling, image vulnerability scanning, and ConfigMap changes without a rollout.
Reference
Change history and rollback
Every change to an application is recorded field by field, deletions included, with a snapshot you can roll back to from the dashboard.
Access control
Passkeys and Google sign-in, roles scoped per area of Telark, and per-action deny rules such as "may edit plans but not approve them".