Skip to content

Kubernetes Probes: Separate Startup, Readiness and Restart Decisions

Define what each Kubernetes probe means for a service so temporary unavailability does not automatically become an unnecessary restart.

GuideDeveloper tools

By

Updated 2 min read
Kubernetes Probe Decisions: cloud illustration with IndieTools branding

Kubernetes startup, readiness and liveness probes answer different questions. Define those questions before pointing all three at a generic endpoint that happens to return a successful status.

The Kubernetes probe guide explains that readiness affects whether a container receives service traffic, while liveness can lead to a restart. Startup probes allow initialization to complete before normal health decisions take over.

Describe the service's useful state

Write down the minimum conditions under which the application can handle its intended requests. Distinguish essential dependencies from optional features.

ASO.dev describes store operations and product research workflows and has a declared Kubernetes association. Its deployment configuration is not disclosed. An operations workspace provides a useful exercise because one external service may be unavailable while other parts of the interface remain useful.

A health policy should reflect that distinction instead of collapsing all degraded behavior into one boolean.

Keep startup separate from recovery

An application may need time to initialize configuration or load required resources. A startup check should represent completion of that initialization.

Do not make normal liveness thresholds so lenient that genuine process failures take too long to detect merely to accommodate startup. Use the appropriate startup mechanism and test its behavior with realistic initialization conditions.

Record the observed startup range rather than choosing a threshold from one unusually fast local run.

Ask whether restarting can help

Liveness should identify a condition for which restarting the container is a sensible recovery action. A temporary outage of a shared external dependency may not meet that criterion.

If every instance restarts because the same remote API is unavailable, the restart cycle may add disruption without restoring the dependency.

Test failure modes deliberately in a controlled environment. Observe both application behavior and the controller's resulting actions.

Make readiness operationally useful

A readiness failure should mean the instance is temporarily unsuitable for its expected traffic. Confirm that the check actually detects the condition it claims to represent.

Also verify recovery. When the dependency or internal state becomes usable again, the instance should return to service according to the intended policy.

Avoid an endpoint that always returns success while the application cannot serve its main operation. Conversely, do not make readiness depend on unrelated optional work.

Verify the user-facing path

A passing probe is not an end-to-end test. Exercise a representative read and a safe write or validation flow in the release environment.

Record probe outcomes, restart counts and request behavior together. This makes it possible to see whether a health configuration improves recovery or merely makes a dashboard green.

The companion Kubernetes scheduled-job guide covers work that may fail independently of web-service health.

Keep probe changes inside the release review. A tiny configuration edit can alter traffic routing and restart behavior across every instance.

The acceptance standard is a documented connection between a failure, its probe result and the resulting operational action. That is more useful than treating every HTTP success as proof that the entire product is healthy. Keep probe endpoints inexpensive so frequent checks do not themselves consume the capacity needed to recover.

More guide articles