RO EN

Resilience in .NET (1/6): why distributed systems fail and what Polly v8 brings

Resilience in .NET (1/6): why distributed systems fail and what Polly v8 brings ✨ Imagine generată cu AI
Doru Bulubașa
30 September 2026
37 views

Last day of the month, 23:40. The billing application sends documents in batches to an external service. The service doesn’t crash — that would have been too easy. It just responds in 40 seconds instead of 400 milliseconds. Ten minutes later, your API also stops responding: the thread pool is full of waiting requests, the liveness probe times out, the orchestrator restarts the pods, and the new pods fill up again in a few seconds.

No bugs in the code. No missed exceptions. Just a slow dependency and an application that wasn’t built to say “I’m not waiting anymore”.

This series is about that: how to build .NET applications that stay standing when the world around them shakes. Six articles, from fundamentals to chaos engineering. In the first, we lay the foundations: what types of failures exist, why latency is more dangerous than errors, and what Polly v8 looks like — the library we use throughout the series.

The network is not reliable. Ever.

In the ’90s, a few engineers at Sun Microsystems formulated what we now call the fallacies of distributed computing: eight assumptions that every programmer makes, consciously or not, and all of which are false. The first three are enough to break an architecture: the network is reliable, latency is zero, bandwidth is infinite.

In a classic monolith, a method call takes nanoseconds and either succeeds or throws a clear exception. When the same call goes over the network — to a database, an external API, another microservice, a distributed cache — a third possibility arises, the most unpleasant: you don’t know. The request may or may not have arrived. It may or may not have been processed. The response may come in 50 ms, 50 seconds, or never.

Resilience doesn’t mean eliminating these situations. It means deciding in advance what your application does when they occur.

Failure taxonomy

The first step is not to install a NuGet package, but to understand what kind of failure you are dealing with. The wrong strategy applied to the wrong type of failure does more harm than no strategy at all.

  • Transient failures — connection reset, 503 during a deploy, 429 from throttling, a SQL deadlock, a cloud database failover. They disappear on their own within a few hundred milliseconds or seconds. Here, retry makes sense.
  • Permanent failures — 400 Bad Request, 401, 404, a changed API contract, invalid data. No matter how many retries you do, the result is the same. Here, retry only wastes resources and delays the error.
  • Latency and degradation — the service is alive but responds ten times slower. You get no error, you just wait. This is the most dangerous type, because nothing “fails” visibly until everything fails.
  • Overload — the dependency is alive but suffocated. Every extra request, including your well-intentioned retries, pushes it deeper. Here you need to reduce the pressure, not increase it: circuit breaker, rate limiting, controlled degradation.

Note that only the first category is solved with retries. For the other three, the reflex “try once more” is exactly the wrong reflex.

Why latency is more dangerous than errors

A service that fails fast is a good neighbor: you get the error in 5 ms, handle it, move on. A service that hangs is a toxic neighbor, and Little’s law explains why.

The number of requests simultaneously in the system equals the arrival rate multiplied by the time each request spends in the system (L = λ × W). At 200 requests per second and 50 ms per call to a dependency, you have on average 10 concurrent requests waiting. If the dependency rises to 30 seconds, you have 6,000.

Each holds allocated memory, a connection from the pool, a socket, maybe an open transaction. The .NET thread pool starts injecting new threads, slowly, while meanwhile requests unrelated to the slow dependency — a simple GET /products — also queue up. This is how a cascading failure is born: a local problem in a single service becomes a global problem across the platform.

That’s why the first resilience strategy you add is not retry, but timeout. Timeout turns unlimited latency into a fast and predictable failure. A fast failure can be handled. An infinite wait cannot.

Polly v8: what changed

Polly has been the de facto standard for resilience in .NET for years. Version 8 was an almost complete rewrite, done in collaboration with the Microsoft .NET team, and the official package Microsoft.Extensions.Http.Resilience is built directly on top of it. If you have code written on Polly 7, the concepts remain, but the API looks different.

  • Policy becomes strategy. Retry, circuit breaker, timeout, fallback, hedging, rate limiter — all are resilience strategies.
  • PolicyWrap becomes ResiliencePipeline. You compose strategies into a pipeline, and the order you add them is the order they execute, from outside to inside.
  • One single API. There is no longer a separation between sync and async policies. Everything is built around ValueTask, with minimal allocations on the happy path.
  • Typed and validated options. Each strategy is configured through an options object (RetryStrategyOptions, CircuitBreakerStrategyOptions, etc.), validated when building the pipeline.
  • Built-in telemetry. Logs and metrics via System.Diagnostics.Metrics, no extra code needed.

The difference is immediately visible. The same retry in Polly 7:

var policy = Policy
    .Handle<HttpRequestException>()
    .WaitAndRetryAsync(3, attempt => TimeSpan.FromSeconds(Math.Pow(2, attempt)));}

And in Polly 8:

var pipeline = new ResiliencePipelineBuilder()
    .AddRetry(new RetryStrategyOptions
    {
        ShouldHandle = new PredicateBuilder().Handle<HttpRequestException>(),
        MaxRetryAttempts = 3,
        BackoffType = DelayBackoffType.Exponential,
        UseJitter = true
    })
    .Build();

More verbose but explicit: each parameter has a name, and jitter — which we discuss in detail in the next article — is a simple property, not a handwritten formula. If you still use Microsoft.Extensions.Http.Polly with AddTransientHttpErrorPolicy, Microsoft recommends migrating to the new resilience package.

First pipeline: a correctly done timeout

The base package is Polly.Core:

dotnet add package Polly.Core
using Polly;
using Polly.Timeout;

ResiliencePipeline pipeline = new ResiliencePipelineBuilder()
    .AddTimeout(TimeSpan.FromSeconds(3))
    .Build();

try
{
    ExchangeRates rates = await pipeline.ExecuteAsync(
        async token => await exchangeRateService.GetRatesAsync(token),
        cancellationToken);
}
catch (TimeoutRejectedException)
{
    logger.LogWarning("The exchange rate service did not respond within 3 seconds.");
}

Seems trivial, but there is a trap almost everyone steps on at the start. Look at this version:

// WRONG: the token received from Polly is ignored
await pipeline.ExecuteAsync(
    async _ => await exchangeRateService.GetRatesAsync(),
    cancellationToken);

In Polly v8, timeout is cooperative. When the 3 seconds expire, Polly cancels the CancellationToken it gave you in the callback — and that’s it. It doesn’t kill threads, it doesn’t abandon the task. If your operation does not propagate the token down to the real I/O (HttpClient, EF Core, Cosmos DB driver, Task.Delay), it continues running, and you get TimeoutRejectedException only after 40 seconds, when the operation finishes on its own. Exactly the problem you wanted to solve.

The golden rule, valid for the whole series: the token received from the pipeline must be passed down to the last asynchronous call. Without this, no timeout, no matter how well configured, protects you.

Pipelines in dependency injection

Building ad-hoc pipelines is useful for experiments. In a real application, you register them in DI, with names, and reuse them. The package is Polly.Extensions:

dotnet add package Polly.Extensions
builder.Services.AddResiliencePipeline("exchange-rate", pipeline =>
{
    pipeline.AddTimeout(TimeSpan.FromSeconds(3));
});

And you consume it via ResiliencePipelineProvider<string>:

public sealed class ExchangeRateReader(
    ResiliencePipelineProvider<string> pipelineProvider,
    IExchangeRateService exchangeRateService)
{
    private readonly ResiliencePipeline _pipeline =
        pipelineProvider.GetPipeline("exchange-rate");

    public async Task<ExchangeRates> ReadAsync(CancellationToken cancellationToken) =>
        await _pipeline.ExecuteAsync(
            async token => await exchangeRateService.GetRatesAsync(token),
            cancellationToken);
}

The benefit is not just aesthetic. Pipelines registered in DI automatically receive ILoggerFactory and metrics, are singletons (the state of a circuit breaker must be shared, otherwise it makes no sense), and can be configured centrally.

For HttpClient: Microsoft.Extensions.Http.Resilience

Most external calls in an ASP.NET Core application go through HttpClient. For these, Microsoft offers a dedicated package that provides a complete pipeline, preconfigured with reasonable defaults:

dotnet add package Microsoft.Extensions.Http.Resilience
builder.Services
    .AddHttpClient<ISupplierClient, SupplierClient>(client =>
    {
        client.BaseAddress = new Uri("https://api.supplier.example/");
    })
    .AddStandardResilienceHandler();

One line, five strategies. It’s worth knowing exactly what you get, in execution order, from outside to inside:

  1. Rate limiter — maximum 1,000 concurrent requests, no queue. Protects the dependency (and your own process) from an uncontrolled surge.
  2. Total request timeout — 30 seconds for the entire operation, including all retries.
  3. Retry — up to 3 retries, exponential backoff with jitter, starting at 2 seconds. Handles 5xx errors, 408, 429, HttpRequestException, and per-attempt timeouts.
  4. Circuit breaker — opens when more than 10% of requests fail, with a minimum of 100 requests in a 30-second window, and stays open for 5 seconds.
  5. Attempt timeout — 10 seconds for each individual attempt.

Think of the pipeline like an onion. The total timeout is outside the retry, so it puts a firm limit on the whole operation, no matter how many retries occur. The per-attempt timeout is inside the retry, so a stuck attempt is cut off after 10 seconds and counted as a failure — both for retry and circuit breaker. If you reversed the order, a single 30-second timeout would swallow all retries, and the circuit breaker would never see individual failures.

Adjusting default values

Defaults are a starting point, not a universal answer. An internal API that usually responds in 80 ms doesn’t need a 10-second per-attempt timeout:

builder.Services
    .AddHttpClient<ISupplierClient, SupplierClient>(client =>
    {
        client.BaseAddress = new Uri("https://api.supplier.example/");
    })
    .AddStandardResilienceHandler(options =>
    {
        options.TotalRequestTimeout.Timeout = TimeSpan.FromSeconds(10);
        options.AttemptTimeout.Timeout = TimeSpan.FromSeconds(2);
        options.Retry.MaxRetryAttempts = 2;
        options.CircuitBreaker.SamplingDuration = TimeSpan.FromSeconds(15);
    });

Options are validated at startup, and two rules will surely catch you at some point: the total timeout must be greater than the per-attempt timeout, and the circuit breaker’s sampling window must be at least twice the per-attempt timeout. If you violate them, the application throws an exception at the first client resolution, not in production at three in the morning — exactly as it should.

Two practical details:

  • HttpClient.Timeout remains active. It defaults to 100 seconds and wraps the entire handler chain. If you set a total timeout greater than 100 seconds, the HttpClient timeout wins. Always keep it above the pipeline’s total timeout.
  • Retry on POST can be dangerous. If the called endpoint is not idempotent, a retry can create a duplicate invoice. For such clients, there is options.Retry.DisableForUnsafeHttpMethods(). We discuss idempotency in detail in the next article.

When the standard pipeline doesn’t fit at all, you build your own with AddResilienceHandler("name", builder => { ... }), using exactly the same strategies as in Polly.Core.

Observability from day one

A resilience strategy you can’t see in action is a strategy you only hope for. Polly v8 emits metrics through a Meter named Polly (retry events, circuit openings, attempt and pipeline durations). If you use OpenTelemetry, it’s enough to add them:

builder.Services.AddOpenTelemetry()
    .WithMetrics(metrics => metrics
        .AddAspNetCoreInstrumentation()
        .AddHttpClientInstrumentation()
        .AddMeter("Polly"));

From here on, a graph of retries per minute tells you more about a dependency’s health than any provider status page. A retry rate that steadily increases is an alarm signal long before users notice anything.

What resilience DOES NOT solve

It’s worth saying clearly before we go further, because Polly is easy to misuse:

  • It doesn’t fix bugs. A NullReferenceException retried three times is still a NullReferenceException, just more expensive.
  • It doesn’t replace capacity. If a dependency is undersized, resilience helps you fail gracefully, not avoid failing.
  • It multiplies across layers. If the gateway retries 3 times, the backend service retries 3 times, and the HTTP client in that service retries 3 times, a single failure can generate up to 64 calls to the final dependency. Resilience must be designed at the system level, not added reflexively in every layer.

Recap

  • Classify the failure before choosing the strategy: transient, permanent, latency, or overload.
  • Every network call has a timeout. No exceptions.
  • Timeout in Polly v8 is cooperative: the CancellationToken must be propagated down to I/O.
  • Register pipelines in DI, don’t build them ad-hoc.
  • For HttpClient, start with AddStandardResilienceHandler() and adjust values to the real dependency profile.
  • Enable Polly metrics from day one.

What’s next in the series

  1. Fundamentals and Polly v8 — this article
  2. Smart retry policies — backoff, jitter, idempotency, and retry storms
  3. Circuit breaker and bulkhead — how to stop a failure from spreading
  4. Advanced health checks — /health/live vs. /health/ready and why the difference matters
  5. Graceful degradation — fallback, cache, hedging, and what the application does when a dependency is down
  6. Chaos engineering basics — how to prove that all the above really works

In the next article, we dive into retry details: why jitter is not optional, what can be safely retried, and what a retry storm that destroys your own infrastructure looks like.