Reference

Insight rules

The 60 setup-review rules Insights evaluates, by family, and the incident kinds it detects.

Every rule is deterministic code in the analyzer; none calls the model. How reviews run and how cards are triaged is explained in Insights.

Incident kinds

The rules pick one incident kind per affected workload, in this order of precedence, with a sub-cause such as image_pull.not_found or probe_failure.readiness.

KindDetected from
oomA container killed for exceeding its memory limit.
image_pullImage pull failures: not found, unauthorized, rate limited, …
crashloopContainers restarting in a loop, including restarts caused by a failing liveness or startup probe.
schedulingPods that cannot be scheduled: insufficient CPU or memory, taints, affinity, topology spread, volumes, host ports.
probe_failureA readiness probe that keeps failing (or a liveness or startup probe before any restart).
resource_pressureEvictions for memory, disk, ephemeral storage or PIDs, and preemption.
rollout_stuckA rollout that stopped progressing.
config_change_regressionHealth degraded shortly after a recorded change.
otherAnything else, such as a container that cannot be created or a volume that cannot mount.

Recommendation rules

A rule fires only when every read it depends on succeeded in full. A rule id is <family>.<rule>. Production-only rules apply when a namespace, or the environment of a plan that covers the workload, matches ANALYZER_PRODUCTION_PATTERN.

Reliability

RuleFlags
reliability.single_replicaA Deployment or StatefulSet runs a single replica.
reliability.no_pdbSeveral replicas but no disruption budget.
reliability.pdb_blocks_evictionA disruption budget that blocks every node drain.
reliability.no_readiness_probeA container behind a Service has no readiness probe.
reliability.no_liveness_probeA container has no liveness probe.
reliability.liveness_same_as_readinessLiveness and readiness use the same check.
reliability.no_startup_probeThe liveness probe restarts a container while it starts.
reliability.replicas_same_nodeAll replicas run on one node.
reliability.rollout_all_at_onceEvery rollout stops all replicas at once.
reliability.short_grace_periodPods are killed without a graceful shutdown.
reliability.revision_history_zeroNo rollout history kept to roll back to.
reliability.deployment_pausedRollouts are paused.
reliability.liveness_single_failureOne failed liveness check restarts the container.
reliability.probe_port_undeclaredA probe targets a port name the container does not declare.
reliability.pdb_blocks_at_min_scaleA disruption budget that blocks drains once the autoscaler scales down to its minimum.

Resources

RuleFlags
resources.no_requestsA container has no CPU or memory requests.
resources.no_memory_limitA container has no memory limit.
resources.limits_without_requestsA limit is set without a request.
resources.memory_near_limitMemory use close to the limit.
resources.cpu_near_limitCPU use close to the limit, so likely throttled.
resources.overprovisionedRequests far above what the container uses.
resources.underprovisionedUse above what the container requests.
resources.oom_historyA recent out-of-memory kill with the current memory limit.

Scaling

RuleFlags
scaling.hpa_min_equals_maxAn autoscaler whose minimum equals its maximum.
scaling.hpa_missing_requestsAn autoscaler that cannot read utilization.
scaling.hpa_at_maxAn autoscaler held at its maximum.
scaling.no_hpa_sustained_loadSustained high CPU use and no autoscaler.
scaling.hpa_inactiveAn autoscaler that cannot compute its metrics, so it does not scale.
scaling.hpa_scale_down_disabledAn autoscaler that never scales down.

Security

RuleFlags
security.privilegedA privileged container.
security.privilege_escalation_allowedPrivilege escalation not forbidden.
security.runs_as_rootA container that runs, or may run, as root.
security.writable_root_fsA writable root filesystem.
security.added_capabilitiesAdded Linux capabilities.
security.host_namespacesThe node's network, PID, or IPC namespace shared.
security.host_pathNode directories mounted.
security.default_service_accountThe namespace's default service account.
security.token_automountAn API token mounted that may not be needed.
security.secrets_in_envSecrets passed as environment variables.
security.plaintext_secret_envSecret-looking literal values in the spec.
security.seccomp_unsetNo seccomp profile.
security.capabilities_not_droppedThe default Linux capabilities kept.
security.host_portPorts bound on the node.
security.run_as_root_groupA container that runs with the root group.
security.proc_mount_unmasked/proc mounted unmasked.

Images

RuleFlags
images.mutable_tagA moving image tag (latest or none).
images.pull_policy_mismatchA moving tag pulled only when missing, so nodes may run different builds.
images.pull_policy_neverAn image never pulled, so pods fail on nodes that lack it.
images.digest_not_pinned_productionA production workload that runs its image by tag, not digest.

Configuration

RuleFlags
config.duplicate_envAn environment variable defined twice.
config.subpath_no_reloadConfiguration mounted with subPath, so later changes never reach the pod.

Networking

RuleFlags
networking.service_selector_mismatchA Service that selects no pods.
networking.service_port_mismatchA Service that targets a port no container declares.
networking.no_network_policyNo network policy selects the pods.
networking.network_policy_allows_allA network policy that admits all traffic.

Change risk

RuleFlags
change_risk.high_velocityAn application that changes very often.
change_risk.frequent_rollbacksTwo or more rollbacks in the last seven days.

Protection

RuleFlags
protection.production_uncoveredA production application with no protection plan.
protection.production_audit_onlyA production application covered only by audit-mode plans.

Consistency

RuleFlags
consistency.image_skewThe same workload runs different images across namespaces.

Not yet available

These checks are catalogued but not shipped, because they would be noisy or need data the analyzer does not read: missing preStop hooks, zone topology spread, measured CPU throttling, vertical autoscaler recommendations, image vulnerability scanning, ConfigMap changes without a rollout, Ingress backends, incidents that follow changes, drift, missing recent snapshots, and namespace-level quotas and priority classes.