Review on-failure restart limits

Bound restart attempts to recovery requirements and monitor repeated failures.

Description

A restart policy helps a service attempt automatic recovery, but an excessive or inconsistent limit can complicate failure handling. on-failure applies when a process exits with a nonzero status; it does not act on unhealthy status alone.

You can bound retries with a value such as on-failure:5. Five is an example limit, not inherently safer than ten for every service. Unlimited or excessive restarts can hide a fault, while too few attempts can stop a service after a transient failure.

Potential impact

  • Repeated restarts can delay investigation of an ongoing fault.
  • Too few attempts can stop a service after a transient error.
  • Inconsistent service policies complicate monitoring and operational response.

Remediation

  • Set a finite retry count, such as restart: on-failure:5, that meets service recovery needs.
  • When using deploy.restart_policy.max_attempts, also review how window determines which failures count. Do not treat the two settings as having identical counting semantics.
  • Keep team standards consistent while accounting for service importance and recovery behavior, and monitor repeated failures.

Examples

The examples compare retry limits. Choose a value appropriate for the service’s actual recovery requirements.

Before

yaml
services:
  customer:
    image: whoa/hello
    restart: on-failure:10

After

yaml
services:
  customer:
    image: whoa/hello
    restart: on-failure:5

Explanation:

  • Before: Up to ten retries are configured after a failed exit. Ten is not inherently unsafe; compare it with actual recovery needs.
  • After: The retry limit is reduced to five. Test that this does not prevent the attempts needed to recover from transient failures.

References