Health checks that tell the truth

ยท 2 min read

A health endpoint that always returns 200 doesn't help anyone. Separate "is the process alive" from "can it do its job", and let App Service act on the answer.

Load balancers, container orchestrators and Azure App Service all decide where to send traffic based on health checks. A good health endpoint lets them stop sending requests to a broken instance. A bad one either says "healthy" while every request fails, or says "unhealthy" over something trivial and takes the whole service down.

Add health checks

builder.Services.AddHealthChecks()
    .AddDbContextCheck<OrdersDb>(tags: ["ready"])
    .AddCheck<ServiceBusHealthCheck>("service-bus", tags: ["ready"]);

var app = builder.Build();

app.MapHealthChecks("/health/live", new HealthCheckOptions { Predicate = _ => false });
app.MapHealthChecks("/health/ready", new HealthCheckOptions { Predicate = c => c.Tags.Contains("ready") });

AddDbContextCheck comes from the Microsoft.Extensions.Diagnostics.HealthChecks.EntityFrameworkCore package and verifies the app can reach its database.

Liveness and readiness are different questions

  • Liveness: is the process working at all? If not, restart it. This check should run no dependencies at all. Above, Predicate = _ => false means it only confirms the app can answer an HTTP request.
  • Readiness: can it serve traffic right now? If not, stop sending it requests, but don't restart it. This is where database and messaging checks belong.

Mixing the two causes outages. If liveness checks the database, a short database blip makes every instance fail liveness and get restarted at once, and the service goes down because of a problem that would have fixed itself.

Writing a custom check

public class ServiceBusHealthCheck(ServiceBusClient client) : IHealthCheck
{
    public async Task<HealthCheckResult> CheckHealthAsync(HealthCheckContext context, CancellationToken ct = default)
    {
        try
        {
            await using var receiver = client.CreateReceiver("orders");
            await receiver.PeekMessageAsync(cancellationToken: ct);
            return HealthCheckResult.Healthy();
        }
        catch (Exception ex)
        {
            return HealthCheckResult.Unhealthy("Can't reach Service Bus", ex);
        }
    }
}

Keep checks cheap and quick. They run every few seconds on every instance, so a slow or expensive check becomes a load problem of its own.

Let App Service use it

App Service has a built-in Health check feature. Set the path, for example /health/ready, under Monitoring > Health check. App Service pings it on every instance, and when one keeps failing, it stops routing traffic to that instance and eventually replaces it. This only helps when you run at least two instances, since with one there's nowhere else to send traffic.

Don't leak details

The default response is just Healthy or Unhealthy. If you add a detailed JSON writer with exception messages for diagnostics, protect that endpoint, or keep the detail in logs instead.

Takeaway

Expose separate liveness and readiness endpoints. Keep liveness free of dependencies, put dependency checks in readiness, keep every check fast, and point App Service's health check at readiness so it can route around a broken instance.