Description
When several Deployment pods share a node, one node failure can significantly affect service availability. Increasing the replica count alone does not ensure high availability.
Pod anti-affinity or topology spread constraints can express distribution requirements. Omitting explicit anti-affinity does not mean that all pods necessarily run on one node.
Potential impact
- A single node failure can affect several replicas at once.
- Overly strict placement requirements can prevent pods from starting when eligible capacity is insufficient.
Remediation
- Configure
pod_anti_affinityor topology spread constraints to meet availability needs. Check selectors and topology labels on nodes. - Preferred rules encourage distribution but do not guarantee it. Review the scheduling constraints of required rules, available capacity and actual pod placement.
Examples
These placement excerpts omit required containers, the Deployment selector and other settings. The after weight of 100 is still a preference; it does not require different nodes.
Before
hcl
resource "kubernetes_deployment" "example" {
metadata {
name = "terraform-example"
labels = {
k8s-app = "prometheus"
}
}
spec {
replicas = 3
}
}
After
hcl
resource "kubernetes_deployment" "example" {
metadata {
name = "terraform-example"
labels = {
k8s-app = "prometheus"
}
}
spec {
replicas = 3
template {
metadata {
labels = {
k8s-app = "prometheus"
}
}
spec {
affinity {
pod_anti_affinity {
preferred_during_scheduling_ignored_during_execution {
weight = 100
pod_affinity_term {
label_selector {
match_labels = {
k8s-app = "prometheus"
}
}
topology_key = "kubernetes.io/hostname"
}
}
}
}
}
}
}
}
Explanation:
- Before: Only the replica count is specified, without explicit anti-affinity. Actual distribution needs checking.
- After: Pods with the same application label are preferably placed on different hostnames. The rule does not automatically relocate existing pods.