Install Telark for production
Choose a size and storage, expose the dashboard over HTTPS, install, and enrol the first admin.
The Quickstart gets Telark running on defaults. This guide covers the choices to make before a production install. Check the system requirements first.
Every setting is a --set key=value flag on the same command. Helm does not remember them, so keep your flags in a values file or script and pass them again on every helm upgrade.
1. Choose a size
app.mode sets replicas, resource requests, API rate limits and disruption budgets for every Telark service.
| You have | Mode |
|---|---|
| An evaluation, a single node, or up to a few hundred applications | minimal |
| Up to about 1,000 applications in production | standard (default) |
| More than 1,000 applications, or a high change rate across many namespaces | performance |
Count roughly one application per Deployment, StatefulSet, DaemonSet or CronJob in the namespaces Telark watches. standard was measured on 1,000 applications: discovery settles around 200 MiB and a full rediscovery pass takes about 4.5 minutes.
You have outgrown a mode when the discovery pod restarts with OOMKilled, the Applications page shows "Discovering applications…" most of the time, or bulk actions such as force sync on many applications crawl.
Add --set vpa.enabled=true in standard or performance to install the Vertical Pod Autoscaler, so requests follow real usage. Its objects appear on the first helm upgrade after the install.
2. Choose storage
standard and performance run two exporter replicas that share the snapshot and report volumes, so they need a ReadWriteMany class:
--set app.persistence.storageClass=<rwx-class>On a single-node cluster, use --set app.singleNode=true instead: one exporter replica on ReadWriteOnce, any class. app.persistence.size (snapshots) and app.persistence.reportsSize (reports) set the volume sizes.
A bound claim cannot change its access mode later. Switching between minimal and the other modes, or toggling app.singleNode, on a live release needs a reinstall; see Upgrade Telark.
3. Choose how to reach the dashboard
The dashboard is a ClusterIP Service on port 8080. The chart ships no ingress or gateway controller; it uses the one you run.
Passkeys only work on a secure origin (https://, or http://localhost) and are bound to the hostname. For anything other than a port-forward, serve HTTPS and pin the relying party to your hostname:
--set app.auth.passkey.id=telark.example.com \
--set app.auth.passkey.origin=https://telark.example.comWith ingress.enabled or gateway.enabled the chart refuses to render until both are set. Pin them yourself for NodePort and LoadBalancer, and terminate TLS in front of those, since they expose plain HTTP.
| Option | Flags |
|---|---|
| Port-forward (evaluation) | none; kubectl port-forward -n telark svc/telark-ui-service 3000:8080 |
| NodePort | services.ui.serviceType=NodePort, optional services.ui.nodePort=30080 |
| LoadBalancer | services.ui.serviceType=LoadBalancer |
| Ingress | ingress.enabled=true, ingress.className, ingress.host, ingress.tls, ingress.annotations |
| Gateway API | gateway.enabled=true, gateway.parentRefs[0].name, gateway.hostnames[0]. Needs Gateway API v1.0+ CRDs and a controller that implements HTTPRoute v1 (Envoy Gateway, NGINX Gateway Fabric, Cilium, Istio). ingress-nginx does not. |
Example: Ingress with a Let's Encrypt certificate
The hostname must already resolve to your ingress controller. A self-signed certificate is not enough: browsers disable WebAuthn on pages with certificate errors.
-
Install ingress-nginx and cert-manager if you do not run them:
helm upgrade --install ingress-nginx ingress-nginx \ --repo https://kubernetes.github.io/ingress-nginx \ --namespace ingress-nginx --create-namespace helm upgrade --install cert-manager cert-manager \ --repo https://charts.jetstack.io \ --namespace cert-manager --create-namespace \ --set crds.enabled=true -
Create a
ClusterIssuer. Replace the email with the address Let's Encrypt should notify:kubectl apply -f - <<'EOF' apiVersion: cert-manager.io/v1 kind: ClusterIssuer metadata: name: letsencrypt spec: acme: server: https://acme-v02.api.letsencrypt.org/directory email: admin@example.com privateKeySecretRef: name: letsencrypt solvers: - http01: ingress: ingressClassName: nginx EOF -
Use these flags in step 4:
--set ingress.enabled=true \ --set ingress.className=nginx \ --set ingress.host=telark.example.com \ --set 'ingress.tls[0].secretName=telark-tls' \ --set 'ingress.tls[0].hosts[0]=telark.example.com' \ --set ingress.annotations."cert-manager\.io/cluster-issuer"=letsencrypt \ --set app.auth.passkey.id=telark.example.com \ --set app.auth.passkey.origin=https://telark.example.com
4. Install
helm install telark oci://ghcr.io/telark/charts/telark -n telark --create-namespace \
--set app.mode=standard \
--set app.persistence.storageClass=<rwx-class> \
--set 'app.auth.bootstrap.admins={you@example.com}' \
<your dashboard flags from step 3>app.auth.bootstrap.admins is required: passkey self-registration is off, so without it nobody could sign in and the chart refuses to render. List several emails if you want more than one admin. These accounts get the Admin role only from a verified identity: a break-glass enrolment, or their first Google sign-in.
The CRDs install with the release as a bundled subchart and are kept on uninstall. If something else manages them, add --set crds.enabled=false.
5. Run one upgrade to enable change authors
Kyverno stamps telark.io/last-modified-by on workloads so change history can name who made each change. That policy renders only once Kyverno's API exists, which is not yet true during the first install. Run the upgrade once, with the same flags:
helm upgrade telark oci://ghcr.io/telark/charts/telark -n telark <same flags as step 4>Changes made before this carry no author. Secrets are never stamped, so history does not name who changed a Secret.
6. Enrol the first admin
kubectl exec -n telark deploy/telark-auth-service -- ./main break-glass --email you@example.com --enrollOpen https://telark.example.com/register?enroll=<token> within 10 minutes and register a passkey.
Other people join in one of three ways: Google sign-in (Set up Google sign-in), an enrolment link an admin creates, or passkey self-registration if you install with --set app.auth.passkey.selfRegistration="true". Self-registered and first-time Google accounts get the ReadOnly role. Leave self-registration off unless only people you trust can reach the dashboard.
7. Verify
kubectl get pods -n telark
helm test telark -n telarkEvery pod should be Ready and the test should pass.
Security defaults
These are on by default. Change them only with a reason.
CRD write guard (app.crdGuard.enabled, app.crdGuard.enforce). A ValidatingAdmissionPolicy rejects writes to every telark.io resource, and to the OIDC trust Secret, from anything but Telark's own service accounts. Edit rights on the telark namespace therefore do not turn into Telark Admin. Allow a break-glass identity with --set 'app.crdGuard.extraAllowedUsers={system:serviceaccount:ops:breakglass}'. enforce=false only audits.
NetworkPolicies (app.networkPolicy.enabled). Default deny for every Telark pod. The service APIs accept traffic only from Telark pods of the same release, the dashboard port is open, and NATS accepts only discovery and the notifier on port 4222. Egress is not restricted. They need a CNI that enforces NetworkPolicy.
NATS users. Discovery publishes with its own user and the notifier consumes with another. The chart generates both passwords.
Kyverno fails open (app.kyverno.failOpen=true). While Kyverno is down or rolling out, the API server admits every request, including ones an enforce plan would block. Workloads stay deployable at the cost of best-effort enforcement. For strict enforcement set both keys; the render fails if they differ:
--set app.kyverno.failOpen=false \
--set kyverno.features.forceFailurePolicyIgnore.enabled=falseFail-closed means an unavailable Kyverno blocks writes to everything its webhooks match, so keep its two admission replicas.
GitOps (cluster-less renders)
The chart generates the service token, both NATS passwords and the OIDC trust Secret, and reads them back on upgrade with lookup. Argo CD, Flux and other helm template pipelines render without a cluster, so every sync would generate new values and break the running pods. Create the Secrets yourself:
kubectl create secret generic telark-service-token -n telark \
--from-literal=token="$(openssl rand -hex 32)"
kubectl create secret generic telark-nats-publisher -n telark \
--from-literal=username=publisher --from-literal=password="$(openssl rand -hex 24)"
kubectl create secret generic telark-nats-consumer -n telark \
--from-literal=username=consumer --from-literal=password="$(openssl rand -hex 24)"
kubectl create secret generic telark-oidc-trust -n telark --from-literal=googleJwkJson=''Then point the chart at them:
--set app.serviceToken.existingSecret=telark-service-token \
--set nats.existingSecrets.publisher=telark-nats-publisher \
--set nats.existingSecrets.consumer=telark-nats-consumer \
--set app.auth.oidc.existingSecret=telark-oidc-trustInsights runtime
Insights runs on Ollama inside the cluster and is on after install, with the granite4:350m model. Setup reviews run on their own, every application every 2 hours. Incident analysis runs when someone chooses Analyze, or automatically once Analyze automatically on incidents and recoveries is turned on in Settings → Local analyzer, where you can also change the model or turn the analyzer off.
| Setup | Flags |
|---|---|
| Connected (default) | none. Ollama downloads the model (about 708 MB) after start and gets HTTPS egress for it. |
| Air-gapped | app.ollama.autoPull=false. No download and no egress; pre-load the model on the volume or in the image. |
| Your own Ollama endpoint | app.ollama.enabled=false, app.ollama.runtimeUrl=http://<host>:11434. A URL only, no key. |
| No Insights | app.ollama.enabled=false, and turn the analyzer off in Settings → Local analyzer. The rest of Telark works without it. |
Autoscaling
In standard and performance, auth, discovery, the notifier and the dashboard scale on CPU from 1 replica up to 3 or 5. The exporter and the analyzer never autoscale. Tune one service:
--set services.discovery.autoscaling.minReplicas=2 \
--set services.discovery.autoscaling.maxReplicas=8 \
--set services.discovery.autoscaling.targetCPUUtilizationPercentage=70Set the same keys under app.serviceDefaults.autoscaling to change every service.
Other settings
metrics-server on kind, minikube or self-signed kubelets
If HPAs show <unknown>, the bundled metrics-server cannot verify kubelet certificates. There, and only there, add:
--set 'metrics-server.args={--kubelet-insecure-tls,--kubelet-preferred-address-types=InternalIP\,ExternalIP\,Hostname}'Existing Kyverno or metrics-server
Skip the bundled copies with --set app.kyverno.enabled=false or --set metrics-server.enabled=false.
Large installs
Redis keeps everything in memory and never evicts. Beyond 2,000 applications, raise redis.master.resources.limits.memory (512 MiB by default) in proportion.
CORS
The dashboard reaches every API through its own proxy on the same origin, so leave services.<svc>.env.CORS_ALLOWED_ORIGINS empty unless a browser app on another origin calls the exporter, discovery, auth or analyzer APIs directly.
Every value, with its default: Helm values.