Description
A restart policy helps a service attempt automatic recovery, but an excessive or inconsistent limit can complicate failure handling. on-failure applies when a process exits with a nonzero status; it does not act on unhealthy status alone.
You can bound retries with a value such as on-failure:5. Five is an example limit, not inherently safer than ten for every service. Unlimited or excessive restarts can hide a fault, while too few attempts can stop a service after a transient failure.
Potential impact
- Repeated restarts can delay investigation of an ongoing fault.
- Too few attempts can stop a service after a transient error.
- Inconsistent service policies complicate monitoring and operational response.
Remediation
- Set a finite retry count, such as
restart: on-failure:5, that meets service recovery needs. - When using
deploy.restart_policy.max_attempts, also review howwindowdetermines which failures count. Do not treat the two settings as having identical counting semantics. - Keep team standards consistent while accounting for service importance and recovery behavior, and monitor repeated failures.
Examples
The examples compare retry limits. Choose a value appropriate for the service’s actual recovery requirements.
Before
yaml
services:
customer:
image: whoa/hello
restart: on-failure:10
After
yaml
services:
customer:
image: whoa/hello
restart: on-failure:5
Explanation:
- Before: Up to ten retries are configured after a failed exit. Ten is not inherently unsafe; compare it with actual recovery needs.
- After: The retry limit is reduced to five. Test that this does not prevent the attempts needed to recover from transient failures.