Health checks

Prerequisites: Container configuration

Health checks tell Kubernetes whether a container is alive and ready to serve traffic. Without them, Kubernetes has no way to know if your application has crashed or is stuck in a bad state. With them, it restarts unhealthy containers and routes traffic only to containers that are confirmed ready.


Why health checks matter

By default, Kubernetes starts routing traffic to your container the moment it starts. But your application might need 30 to 60 seconds to initialize: connecting to the database, loading configuration, warming caches. Without a health check, traffic arrives before the app is ready and those early requests fail.

Health checks close that gap. Kubernetes waits until your container reports healthy before sending it traffic, and restarts it if it stops responding. During rolling updates, traffic moves to new containers only after they pass their checks.


Health check types

HTTP

The most common type for web applications. Kubernetes sends an HTTP GET request to a specified path and considers the container healthy if it gets a 2xx response.

x-clouve-healthcheck:
  enabled: true
  type: HTTP
  path: /health
  port: 8080
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Use HTTP checks for web apps, REST APIs, GraphQL servers, and any other container with an HTTP endpoint.

The /health endpoint should:

  • Return 200 OK when the app is healthy
  • Return a non-2xx status, or time out, when it is not
  • Respond quickly, in under 1 second
  • Work without authentication
  • Optionally check internal dependencies such as database connectivity

A minimal Express.js health endpoint:

app.get("/health", (req, res) => {
  res.status(200).json({ status: "ok" });
});

TCP

Kubernetes opens a TCP connection to the specified port and treats the container as healthy if the connection succeeds. The check verifies only that the port is listening; there is no application-level checking.

x-clouve-healthcheck:
  enabled: true
  type: TCP
  path: ""
  port: 5432
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Use TCP checks for databases (PostgreSQL, MySQL, MariaDB), caches (Redis, Memcached), message brokers, and any container where you can't make an HTTP call but can verify that a port is listening.

For TCP checks, the path field must be an empty string "".

Command

Kubernetes runs a command inside the container. Exit code 0 means healthy; any other exit code means unhealthy.

x-clouve-healthcheck:
  enabled: true
  type: Command
  path: "pg_isready -U postgres"
  port: 0
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Use Command checks when a custom command gives a better signal than a TCP connection, or when you want to check application state with a CLI tool bundled in the image.

For Command checks, the path field contains the command string to execute. The port field is ignored.


Configuration fields

FieldTypeDescriptionDefault
enabledbooleanWhether health checking is activefalse
typestringHTTP, TCP, or CommandHTTP
pathstringHTTP path or command string (empty for TCP)/health
portnumberPort to check; 0 uses the container's declared port0
initialDelaynumberSeconds to wait before the first health check30
intervalnumberSeconds between health checks10
timeoutnumberSeconds to wait for a response before failing5
failureThresholdnumberConsecutive failures before marking unhealthy3
successThresholdnumberConsecutive successes before marking healthy1

Recommended settings by container type

Web application (HTTP)

x-clouve-healthcheck:
  enabled: true
  type: HTTP
  path: /health
  port: 0 # Uses container port
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Applications with long startup (Java, Moodle, Odoo)

These applications can take 2 to 5 minutes to fully initialize. Increase initialDelay so the container isn't restarted before it finishes starting:

x-clouve-healthcheck:
  enabled: true
  type: HTTP
  path: /
  port: 80
  initialDelay: 120 # 2 minutes - allows full initialization
  interval: 15
  timeout: 10
  failureThreshold: 5
  successThreshold: 1

PostgreSQL database

x-clouve-healthcheck:
  enabled: true
  type: TCP
  path: ""
  port: 5432
  initialDelay: 20
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Or use the pg_isready command for a more accurate check:

x-clouve-healthcheck:
  enabled: true
  type: Command
  path: "pg_isready -U postgres"
  port: 0
  initialDelay: 20
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

MySQL / MariaDB

x-clouve-healthcheck:
  enabled: true
  type: TCP
  path: ""
  port: 3306
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

Redis cache

x-clouve-healthcheck:
  enabled: true
  type: TCP
  path: ""
  port: 6379
  initialDelay: 10
  interval: 5
  timeout: 3
  failureThreshold: 3
  successThreshold: 1

Nginx static site

x-clouve-healthcheck:
  enabled: true
  type: HTTP
  path: /
  port: 80
  initialDelay: 10
  interval: 5
  timeout: 3
  failureThreshold: 2
  successThreshold: 1

Disabling health checks

To disable health checks, set enabled: false. All other fields must still be present:

x-clouve-healthcheck:
  enabled: false
  type: HTTP
  path: /health
  port: 80
  initialDelay: 30
  interval: 10
  timeout: 5
  failureThreshold: 3
  successThreshold: 1

The fields stay so the configuration round-trips: if you re-enable health checks later by setting enabled: true, all your settings are preserved.

Disabling makes sense for simple utility containers with no testable endpoint, for development and test scenarios, and for containers where health checking is impractical.


Native Docker healthcheck emission

When you download a configuration (or upload one and re-download it), each enabled health check is also written as a standard Docker Compose healthcheck: block on the service, alongside the x-clouve-healthcheck extension that the platform reads. The same manifest then works under plain docker-compose up.

Clouve typeGenerated Docker test command
HTTP["CMD", "wget", "--spider", "-q", "http://localhost:<port><path>"]
TCP["CMD", "nc", "-z", "localhost", "<port>"]
Command["CMD-SHELL", "<command from path>"]

interval, timeout, retries (=failureThreshold), and start_period (=initialDelay) are mirrored from the x-clouve-healthcheck values. When enabled: false, no Docker healthcheck: block is emitted.

Tip: For local testing, your image must contain the relevant probe binary: wget for HTTP, nc for TCP. Most Alpine/Debian images include them; minimal scratch/distroless images may not. The Kubernetes-side probe, which is what the platform actually runs in production, does not depend on these binaries.


Timing parameters

Container starts
       │
       │  initialDelay: 30s
       │  (Kubernetes waits before first check)
       │
       ▼
First health check
       │
       │  interval: 10s
       │  (Time between checks)
       │
       ▼
       ... checks every 10s ...
       │
       │  If check fails:
       │  - Wait for next interval
       │  - Try again
       │  - After failureThreshold (3) consecutive failures → UNHEALTHY → restart
       │
       │  If container was unhealthy and check succeeds:
       │  - Need successThreshold (1) consecutive successes → HEALTHY → receive traffic

Sizing guidance:

  • initialDelay: Set this to slightly longer than your application's typical cold-start time; your logs will show how long initialization takes. Common values are 10 to 30 seconds for fast apps and 60 to 120 seconds for heavy apps (Moodle, Odoo).
  • interval: How often to check. Shorter intervals detect failures faster but add overhead. 10 seconds is a good default.
  • timeout: Your health endpoint should respond in well under 1 second, so timeout: 5 is generous.
  • failureThreshold: How many consecutive failures before the container counts as unhealthy. A value of 3 keeps a single transient failure from triggering a restart.

Troubleshooting health check failures

Container keeps restarting (CrashLoopBackOff)

The health check is failing repeatedly and Kubernetes is restarting the container. Common causes:

  1. Wrong path: the /health endpoint doesn't exist. Check the actual path in your app.
  2. Wrong port: the health check is hitting a port your app doesn't listen on. Verify the port.
  3. Initial delay too short: the app isn't ready when the first check runs. Increase initialDelay.
  4. App crashes on startup: the health check isn't the problem; the app itself is failing. Check the application logs.

Health check never succeeds after deploy

  1. Verify the container port matches the health check port.
  2. Verify the health endpoint returns a 2xx status code.
  3. Verify the endpoint doesn't require authentication.
  4. Increase timeout if the response is slow.

Health check succeeds but app isn't working

Your health endpoint isn't checking enough. Make it verify critical dependencies:

app.get("/health", async (req, res) => {
  try {
    // Check database connectivity
    await db.query("SELECT 1");
    res.status(200).json({ status: "ok" });
  } catch (err) {
    res.status(503).json({ status: "error", message: err.message });
  }
});

Next steps