Review the need for container liveness probes

Use liveness probes for unhealthy states that a restart can resolve.

Description

A running container process can still leave an application unresponsive through deadlock or indefinite waiting. A liveness probe checks for such states and helps the kubelet restart the container.

It is not mandatory for every workload. It is separate from restart_policy handling of process exit, and an incorrect probe can cause repeated restarts or cascading failures.

Potential impact

  • A failure that restarting could resolve may persist until manual intervention.
  • Treating a shared dependency outage as liveness failure can unnecessarily restart many instances.

Remediation

  • Use a supported HTTP, TCP or exec liveness_probe to check states that require restarting. Do not make failure depend solely on the availability of an external service.
  • Tune periods and thresholds for actual startup and recovery behavior, and use a startup probe where needed. Use readiness separately to determine whether traffic can be served.

Examples

The after example assumes an application with a separately implemented /health endpoint. The default nginx image does not provide it automatically; configure the endpoint and use a maintained image.

Before

hcl
resource "kubernetes_pod" "app" {
  metadata {
    name = "app"
  }

  spec {
    container {
      name  = "app"
      image = "nginx:1.27"

      port {
        container_port = 80
      }
    }
  }
}

After

hcl
resource "kubernetes_pod" "app" {
  metadata {
    name = "app"
  }

  spec {
    container {
      name  = "app"
      image = "nginx:1.27"

      port {
        container_port = 80
      }

      liveness_probe {
        http_get {
          path = "/health"
          port = 80
        }

        initial_delay_seconds = 10
        period_seconds        = 10
      }
    }
  }
}

Explanation:

  • Before: There is no liveness check for unresponsiveness. restart_policy still separately handles process exit.
  • After: Persistent HTTP probe failure triggers a container restart. The thresholds and endpoint must match actual recovery criteria.

References