Reference
Insight rules
The 60 setup-review rules Insights evaluates, by family, and the incident kinds it detects.
Every rule is deterministic code in the analyzer; none calls the model.
How reviews run and how cards are triaged is explained in
Insights.
The rules pick one incident kind per affected workload, in this order of
precedence, with a sub-cause such as image_pull.not_found or
probe_failure.readiness.
| Kind | Detected from |
|---|
oom | A container killed for exceeding its memory limit. |
image_pull | Image pull failures: not found, unauthorized, rate limited, … |
crashloop | Containers restarting in a loop, including restarts caused by a failing liveness or startup probe. |
scheduling | Pods that cannot be scheduled: insufficient CPU or memory, taints, affinity, topology spread, volumes, host ports. |
probe_failure | A readiness probe that keeps failing (or a liveness or startup probe before any restart). |
resource_pressure | Evictions for memory, disk, ephemeral storage or PIDs, and preemption. |
rollout_stuck | A rollout that stopped progressing. |
config_change_regression | Health degraded shortly after a recorded change. |
other | Anything else, such as a container that cannot be created or a volume that cannot mount. |
A rule fires only when every read it depends on succeeded in full. A
rule id is <family>.<rule>. Production-only rules apply when a
namespace, or the environment of a plan that covers the workload,
matches ANALYZER_PRODUCTION_PATTERN.
| Rule | Flags |
|---|
reliability.single_replica | A Deployment or StatefulSet runs a single replica. |
reliability.no_pdb | Several replicas but no disruption budget. |
reliability.pdb_blocks_eviction | A disruption budget that blocks every node drain. |
reliability.no_readiness_probe | A container behind a Service has no readiness probe. |
reliability.no_liveness_probe | A container has no liveness probe. |
reliability.liveness_same_as_readiness | Liveness and readiness use the same check. |
reliability.no_startup_probe | The liveness probe restarts a container while it starts. |
reliability.replicas_same_node | All replicas run on one node. |
reliability.rollout_all_at_once | Every rollout stops all replicas at once. |
reliability.short_grace_period | Pods are killed without a graceful shutdown. |
reliability.revision_history_zero | No rollout history kept to roll back to. |
reliability.deployment_paused | Rollouts are paused. |
reliability.liveness_single_failure | One failed liveness check restarts the container. |
reliability.probe_port_undeclared | A probe targets a port name the container does not declare. |
reliability.pdb_blocks_at_min_scale | A disruption budget that blocks drains once the autoscaler scales down to its minimum. |
| Rule | Flags |
|---|
resources.no_requests | A container has no CPU or memory requests. |
resources.no_memory_limit | A container has no memory limit. |
resources.limits_without_requests | A limit is set without a request. |
resources.memory_near_limit | Memory use close to the limit. |
resources.cpu_near_limit | CPU use close to the limit, so likely throttled. |
resources.overprovisioned | Requests far above what the container uses. |
resources.underprovisioned | Use above what the container requests. |
resources.oom_history | A recent out-of-memory kill with the current memory limit. |
| Rule | Flags |
|---|
scaling.hpa_min_equals_max | An autoscaler whose minimum equals its maximum. |
scaling.hpa_missing_requests | An autoscaler that cannot read utilization. |
scaling.hpa_at_max | An autoscaler held at its maximum. |
scaling.no_hpa_sustained_load | Sustained high CPU use and no autoscaler. |
scaling.hpa_inactive | An autoscaler that cannot compute its metrics, so it does not scale. |
scaling.hpa_scale_down_disabled | An autoscaler that never scales down. |
| Rule | Flags |
|---|
security.privileged | A privileged container. |
security.privilege_escalation_allowed | Privilege escalation not forbidden. |
security.runs_as_root | A container that runs, or may run, as root. |
security.writable_root_fs | A writable root filesystem. |
security.added_capabilities | Added Linux capabilities. |
security.host_namespaces | The node's network, PID, or IPC namespace shared. |
security.host_path | Node directories mounted. |
security.default_service_account | The namespace's default service account. |
security.token_automount | An API token mounted that may not be needed. |
security.secrets_in_env | Secrets passed as environment variables. |
security.plaintext_secret_env | Secret-looking literal values in the spec. |
security.seccomp_unset | No seccomp profile. |
security.capabilities_not_dropped | The default Linux capabilities kept. |
security.host_port | Ports bound on the node. |
security.run_as_root_group | A container that runs with the root group. |
security.proc_mount_unmasked | /proc mounted unmasked. |
| Rule | Flags |
|---|
images.mutable_tag | A moving image tag (latest or none). |
images.pull_policy_mismatch | A moving tag pulled only when missing, so nodes may run different builds. |
images.pull_policy_never | An image never pulled, so pods fail on nodes that lack it. |
images.digest_not_pinned_production | A production workload that runs its image by tag, not digest. |
| Rule | Flags |
|---|
config.duplicate_env | An environment variable defined twice. |
config.subpath_no_reload | Configuration mounted with subPath, so later changes never reach the pod. |
| Rule | Flags |
|---|
networking.service_selector_mismatch | A Service that selects no pods. |
networking.service_port_mismatch | A Service that targets a port no container declares. |
networking.no_network_policy | No network policy selects the pods. |
networking.network_policy_allows_all | A network policy that admits all traffic. |
| Rule | Flags |
|---|
change_risk.high_velocity | An application that changes very often. |
change_risk.frequent_rollbacks | Two or more rollbacks in the last seven days. |
| Rule | Flags |
|---|
protection.production_uncovered | A production application with no protection plan. |
protection.production_audit_only | A production application covered only by audit-mode plans. |
| Rule | Flags |
|---|
consistency.image_skew | The same workload runs different images across namespaces. |
These checks are catalogued but not shipped, because they would be noisy
or need data the analyzer does not read: missing preStop hooks, zone
topology spread, measured CPU throttling, vertical autoscaler
recommendations, image vulnerability scanning, ConfigMap changes without
a rollout, Ingress backends, incidents that follow changes, drift,
missing recent snapshots, and namespace-level quotas and priority
classes.