Tuesday, 14:07. The database has an I/O problem and responds in 3 seconds to any query, for about a minute. Unpleasant, but not dramatic: the application would have responded more slowly for a minute, then everything would have returned to normal.
But the health check endpoint also checks the database, and the same endpoint is used as a liveness probe in Kubernetes. The probe has a one-second timeout. Three consecutive failed checks, and the kubelet decides the application is dead and restarts it. On all 8 pods, almost simultaneously.
The new pods start cold: without cache, without connections in the pool, with the JIT still working. All try to open connections to a database that is struggling anyway. The health check fails again. Restart again. At 14:30 the I/O problem was long resolved, but the application is still in a restart loop it caused itself.
Health checks are the mechanism by which the application communicates its state to the surrounding infrastructure: the orchestrator, the load balancer, the monitoring system. If you configure them incorrectly, not only do they not protect you, but they become the cause of the incident. In this article we see what a correct design looks like.
Three different questions
The mistake in the story above comes from confusion: a single /health endpoint that tries to answer three different questions. Kubernetes, and generally any modern orchestrator, separates them, because each answer triggers a different action:
- Startup: Have you finished starting up? As long as the answer is “no”, the other two probes are suspended. If startup takes too long, the container is restarted.
- Liveness: Are you still alive? Would a restart help? A repeated failure means restart. It is the most drastic action, so the question must be the narrowest.
- Readiness: Can you receive traffic now? A failure means the pod is removed from the load balancer, without restart. When the answer becomes positive again, traffic returns.
From this results the basic rule of the entire article:
Liveness checks only problems that a restart solves. Readiness checks only problems that waiting solves. What neither solves has no place in probes.
A slow database is not fixed by restarting the application. So it has no place in liveness.
Health checks in ASP.NET Core: the basics
ASP.NET Core has native support for health checks. Each check implements IHealthCheck and returns one of three results: Healthy, Degraded, or Unhealthy. By default, the endpoint responds with 200 for Healthy and Degraded and with 503 for Unhealthy.
The key to the design is tags. You register all checks once, tag them by role, then expose different endpoints, each with its own filter:
builder.Services.AddHealthChecks()
.AddCheck<StartupHealthCheck>("startup", tags: ["startup", "ready"])
.AddCheck<OutboxWorkerHealthCheck>("outbox-worker", tags: ["live"])
.AddDbContextCheck<AppDbContext>("database",
failureStatus: HealthStatus.Degraded, tags: ["deps"]);
var app = builder.Build();
app.MapHealthChecks("/health/startup", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("startup")
});
app.MapHealthChecks("/health/live", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("live")
});
app.MapHealthChecks("/health/ready", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("ready")
});
AddDbContextCheck comes from the Microsoft.Extensions.Diagnostics.HealthChecks.EntityFrameworkCore package and checks if a connection can be opened through the EF Core context. For other dependencies (Redis, Cosmos DB, Azure Service Bus, RabbitMQ), the community packages AspNetCore.HealthChecks.* have ready-made checks.
Note that the database has the tag deps, not live or ready. We will return immediately to the reason.
Liveness: as narrow as possible
The safest variant of a liveness probe is the one that checks nothing. If the process receives the HTTP request and responds, it is alive. For most applications this is sufficient:
app.MapHealthChecks("/health/live", new HealthCheckOptions
{
Predicate = _ => false // no checks: just "process responds"
});
With Predicate = _ => false, the endpoint returns 200 Healthy without running any checks. If the process is blocked, the thread pool is exhausted, or the application has entered a state where it can no longer respond to HTTP, the probe times out and Kubernetes does exactly what it should: restart.
When is it worth adding something to liveness? Only for internal states from which the application cannot recover by itself, but from which a restart definitely helps. The classic example is a background worker that has blocked: a queue consumer or an Outbox processor that no longer processes anything, although the process responds perfectly to HTTP.
public sealed class WorkerHeartbeat
{
private long _lastBeatTicks = DateTime.UtcNow.Ticks;
public void Beat() =>
Interlocked.Exchange(ref _lastBeatTicks, DateTime.UtcNow.Ticks);
public TimeSpan SinceLastBeat =>
DateTime.UtcNow - new DateTime(Interlocked.Read(ref _lastBeatTicks), DateTimeKind.Utc);
}
public sealed class OutboxWorkerHealthCheck(WorkerHeartbeat heartbeat) : IHealthCheck
{
private static readonly TimeSpan MaxSilence = TimeSpan.FromMinutes(2);
public Task<HealthCheckResult> CheckHealthAsync(
HealthCheckContext context, CancellationToken cancellationToken = default)
{
var silence = heartbeat.SinceLastBeat;
return Task.FromResult(silence < MaxSilence
? HealthCheckResult.Healthy()
: HealthCheckResult.Unhealthy($"The Outbox worker has not reported for {silence.TotalSeconds:N0}s."));
}
}
The worker calls heartbeat.Beat() at each iteration of the loop, including when it has nothing to process. Otherwise, an empty queue would look exactly like a blocked worker. WorkerHeartbeat is registered as a singleton, so the worker and the health check see the same instance.
Startup: for applications that start slowly
Some applications need time before they are usable: they apply migrations, load a reference cache, warm up connections, download configurations. Without a startup probe, the only solution is to guess an initialDelaySeconds for liveness. If it is too small, Kubernetes kills the application before it finishes starting. If it is too large, a container stuck at startup is detected very late.
The startup probe elegantly solves the problem. ASP.NET Core does not have a predefined startup check, but the pattern is simple: a singleton with a flag that a background service sets when it has finished.
public sealed class StartupHealthCheck : IHealthCheck
{
private volatile bool _completed;
public void MarkCompleted() => _completed = true;
public Task<HealthCheckResult> CheckHealthAsync(
HealthCheckContext context, CancellationToken cancellationToken = default) =>
Task.FromResult(_completed
? HealthCheckResult.Healthy("Startup completed.")
: HealthCheckResult.Unhealthy("Startup in progress."));
}
public sealed class WarmupService(
StartupHealthCheck startup,
IServiceScopeFactory scopeFactory) : BackgroundService
{
protected override async Task ExecuteAsync(CancellationToken stoppingToken)
{
await using var scope = scopeFactory.CreateAsyncScope();
var cache = scope.ServiceProvider.GetRequiredService<ReferenceDataCache>();
await cache.LoadAsync(stoppingToken); // reference data, exchange rates, settings
startup.MarkCompleted();
}
}
builder.Services.AddSingleton<StartupHealthCheck>();
builder.Services.AddHostedService<WarmupService>();
AddCheck<StartupHealthCheck> uses the singleton instance from the container, so the health check and the warm-up service work with the same object. In the configuration above, the same check also has the ready tag: as long as the warm-up is not finished, the pod does not receive traffic.
Readiness: shared dependencies are a trap
Readiness seems the natural place for the database: if the database is down, the pod cannot serve requests, so it is not “ready”. Logical, but dangerous.
Think about what happens when the database has a problem. It is not just one pod, but all. All declare readiness failure at the same time, and Kubernetes removes all from the load balancer. The ingress no longer has any backend and responds with 503 to absolutely everything, including:
- static pages and endpoints that do not need the database;
- responses that could have been served from cache;
- your nice message like “Billing is temporarily unavailable, please try again in a few minutes”;
- administration endpoints with which you could have diagnosed the problem.
You have turned a degraded functionality into a completely down application. The exact same trap was signaled in the article about circuit breaker, related to the breaker state.
Readiness must answer the question “is this instance able to serve traffic?”, not “is the whole system healthy?”. Things that belong to the instance have their place here: unfinished warm-up, a local cache that has not loaded, an overloaded instance that wants to receive less traffic. A dependency shared by all instances has no place here, because removing the pod from traffic does not fix it and does not move traffic to a healthier instance. All are equally affected.
There is a legitimate exception: per-instance dependencies, for example a local sidecar or a connection that can only block on a certain pod. There, removing the pod from traffic really helps, because the other instances take over the requests.
Dependencies: a separate endpoint, for people
This does not mean that the state of the database or external APIs does not matter. It matters enormously, but for monitoring and people, not for the orchestrator. You put them on a separate endpoint, with a detailed response, and let it return Degraded instead of Unhealthy when a dependency is down:
app.MapHealthChecks("/health/deps", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("deps"),
ResponseWriter = WriteJsonResponse
})
.RequireAuthorization("Operations");
The default response is simple text (Healthy, Degraded). For people and dashboards, a detailed JSON is much more useful:
static Task WriteJsonResponse(HttpContext context, HealthReport report) =>
context.Response.WriteAsJsonAsync(new
{
status = report.Status.ToString(),
totalDurationMs = report.TotalDuration.TotalMilliseconds,
checks = report.Entries.Select(entry => new
{
name = entry.Key,
status = entry.Value.Status.ToString(),
description = entry.Value.Description,
durationMs = entry.Value.Duration.TotalMilliseconds,
tags = entry.Value.Tags
})
});
Two security notes. First thing: the detailed endpoint must not be public. It shows you the names of dependencies, topology, and response times, exactly the map someone who wants to attack the system needs. Protect it with authorization or expose it only on an internal port with RequireHost("*:8081"). Second thing: do not put exception messages in description. A connection string appearing in an error message is a classic data leak.
Circuit breaker as a health check
The circuit breaker state from the previous article is an excellent indicator for the dependencies endpoint. CircuitBreakerStateProvider from Polly v8 gives you the state at any moment:
var anafCircuit = new CircuitBreakerStateProvider();
builder.Services.AddResiliencePipeline("anaf", pipeline =>
pipeline.AddCircuitBreaker(new CircuitBreakerStrategyOptions
{
StateProvider = anafCircuit
// ... rest of configuration
}));
builder.Services.AddHealthChecks()
.AddCheck("anaf-circuit", () => anafCircuit.CircuitState switch
{
CircuitState.Closed => HealthCheckResult.Healthy(),
CircuitState.HalfOpen => HealthCheckResult.Degraded("Circuit under test."),
_ => HealthCheckResult.Degraded("Circuit open: ANAF service not responding.")
}, tags: ["deps"]);
The result is Degraded, not Unhealthy. The application works, only a functionality is temporarily unavailable. This is exactly the nuance the Degraded state was created to express.
Cheap and fast health checks
Health checks run frequently. With 10 pods, probes every 5–10 seconds, plus the load balancer and monitoring system, you easily reach hundreds of checks per minute. A few rules:
- Minimal checks. Opening a connection or a
SELECT 1, not a query on the invoices table. A health check is not a functional test. - Do not call external APIs from probes. Besides adding latency, you consume the provider’s rate limit quota just to find out what the circuit breaker already tells you.
- Explicit timeout per check. In Kubernetes,
timeoutSecondsdefaults to 1 second. A check that takes 3 seconds does not return “slow”, but a probe failure.AddCheckaccepts atimeoutparameter; use it to stay under the probe timeout.
Health check publisher
For monitoring, you don’t need to wait for someone to query the endpoint. IHealthCheckPublisher runs checks periodically in the background and sends you the report, which you can write to logs, metrics, or Application Insights:
builder.Services.Configure<HealthCheckPublisherOptions>(options =>
{
options.Delay = TimeSpan.FromSeconds(10);
options.Period = TimeSpan.FromSeconds(30);
options.Predicate = check => check.Tags.Contains("deps");
});
builder.Services.AddSingleton<IHealthCheckPublisher, LoggingHealthCheckPublisher>();
public sealed class LoggingHealthCheckPublisher(
ILogger<LoggingHealthCheckPublisher> logger) : IHealthCheckPublisher
{
public Task PublishAsync(HealthReport report, CancellationToken cancellationToken)
{
foreach (var (name, entry) in report.Entries.Where(e => e.Value.Status != HealthStatus.Healthy))
{
logger.LogWarning("Health check {Name}: {Status} ({Description})",
name, entry.Status, entry.Description);
}
return Task.CompletedTask;
}
}
This way you get alerts like “database is Degraded for 3 minutes” without the orchestrator making any decisions based on this information. People decide, not the kubelet.
Configuration in Kubernetes
With separate endpoints, probe configuration becomes clear. .NET 8+ images listen by default on port 8080:
containers:
- name: api
image: registry.example/facturare-api:1.4.2
ports:
- containerPort: 8080
startupProbe:
httpGet:
path: /health/startup
port: 8080
periodSeconds: 2
failureThreshold: 30 # up to 60s for startup
livenessProbe:
httpGet:
path: /health/live
port: 8080
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3 # restart after ~30s without response
readinessProbe:
httpGet:
path: /health/ready
port: 8080
periodSeconds: 5
timeoutSeconds: 2
failureThreshold: 2
Note the asymmetry: liveness is patient (30 seconds without response before restart), because a restart has a high cost. Readiness reacts faster, because temporary removal from traffic is cheap and reversible.
Graceful shutdown
Probes have a lesser-known problem at shutdown too. When Kubernetes stops a pod, during a deploy or scale-down, it sends SIGTERM to the container and, in parallel, starts removing the pod from the service’s endpoint list. The two happen simultaneously, not in order. For a few seconds, the load balancer may still send requests to a pod that has already started shutting down, and these fail.
The usual solution is a preStop hook that delays SIGTERM for a few seconds, during which the removal from traffic propagates:
lifecycle:
preStop:
exec:
command: ["sleep", "5"]
Warning: chiseled or distroless images do not have sleep. Recent Kubernetes versions have a native sleep action for lifecycle hooks; check if your cluster supports it. After SIGTERM, ASP.NET Core finishes ongoing requests within the HostOptions.ShutdownTimeout limit, which must fit, together with preStop, into terminationGracePeriodSeconds (default 30 seconds).
Anti-patterns, briefly
- A single
/healthendpoint used for all three probes. - Database or any shared dependency in liveness.
- Shared dependencies in readiness that remove all pods from traffic simultaneously.
- Calls to external APIs on every probe.
- Lack of a startup probe, compensated with a guessed
initialDelaySeconds. - Checks slower than the probe’s
timeoutSeconds. - Detailed endpoint exposed publicly, with exception messages in the response.
Summary
- Startup, liveness, and readiness are three different questions, with three different consequences.
- Liveness checks only what a restart solves. The safest variant: just “process responds”, plus heartbeats for background workers.
- Readiness checks only the instance’s state, not shared dependencies.
- Dependencies go on a separate, protected endpoint, with
Degradedresult and detailed JSON response. - Tags allow a single registration of checks and multiple filtered endpoints.
IHealthCheckPublisherfor continuous monitoring, independent of probes.- Cheap probes, with explicit timeouts, plus
preStopfor graceful shutdown.
Resilience in .NET series
- Fundamentals and Polly v8
- Smart retry policies
- Circuit breaker and bulkhead
- Advanced health checks — this article
- Graceful degradation
- Chaos engineering basics
We repeated the same promise in the last three articles: “what do you do when a dependency is down?”. In the next article we finally answer: graceful degradation, with fallback, serving data from cache, hedging, feature flags, and a concrete example of what a billing application does when the e-Invoice system does not respond.