RO EN

Resilience in .NET (5/6): graceful degradation, or what to do when a dependency is down

Resilience in .NET (5/6): graceful degradation, or what to do when a dependency is down ✨ Imagine generată cu AI
Doru Bulubașa
09 October 2026
35 views

The last working day of the month, 3:30 PM. The national e-Factura system no longer responds: unannounced maintenance or overload, it doesn’t matter. Two companies use two different invoicing applications.

In the first application, the Issue Invoice button shows a spinner for 30 seconds, then a red error: “The ANAF service is not responding. Please try again later.” The accountant tries a few more times, then calls support. The month’s invoices remain unissued until the evening.

In the second application, the button responds instantly. The invoice is issued, numbered, and saved, and next to it appears a small yellow badge: “Waiting to transmit to ANAF.” At the top, a discreet banner: “Transmission to SPV is temporarily delayed. Invoices are saved and will be transmitted automatically.” At 6:40 PM, when ANAF comes back online, all invoices are sent automatically. The accountant finished work at 4:00 PM.

Both applications had the same failed dependency. Only one of them was designed to work with a failed dependency. This is what graceful degradation is about: not how to avoid failures, which we covered in previous articles, but what you offer the user when, inevitably, they occur.

All roads led here

In the last three articles, we always postponed the same question. Retry stops after a number of attempts. Circuit breaker throws BrokenCircuitException in microseconds. Health check reports Degraded. All these mechanisms give you a fast and predictable error. None tells you what to do with it.

The answer is not technical but product-related. And it starts with a question you must ask for each functionality: what is the most useful thing I can offer the user if dependency X is missing?

Classify functionalities before writing code

Controlled degradation is designed, not improvised during an incident. The first step is a simple classification, which you do together with the product owner:

  • Critical — without them, the application makes no sense: issuing invoices, authentication, payment. For these, you look for alternative ways, not giving up.
  • Important — useful but can work with older data or in a simplified form: exchange rates, statistics dashboard, advanced search.
  • Optional — can temporarily disappear without the user losing anything essential: recommendations, real-time notifications, avatars, advanced PDF export.

For each functionality, write on one line: what it depends on and what happens when that dependency fails. The resulting table becomes the reference document for the rest of the article. Each strategy below corresponds to a line in it.

Strategy 1: fallback with Polly v8

The basic mechanism is the fallback strategy: when the pipeline fails, instead of propagating the exception, you return an alternative result. For optional functionalities, the alternative result is often simply “nothing”:

builder.Services.AddResiliencePipeline<string, IReadOnlyList<Produs>>("recommendations", pipeline =>
{
    pipeline
        .AddFallback(new FallbackStrategyOptions<IReadOnlyList<Produs>>
        {
            ShouldHandle = new PredicateBuilder<IReadOnlyList<Produs>>()
                .Handle<HttpRequestException>()
                .Handle<TimeoutRejectedException>()
                .Handle<BrokenCircuitException>(),
            FallbackAction = _ => Outcome.FromResultAsValueTask<IReadOnlyList<Produs>>([])
        })
        .AddRetry(retryOptions)
        .AddCircuitBreaker(breakerOptions)
        .AddTimeout(TimeSpan.FromSeconds(2));
});

The fallback is always outside the pipeline. You want retry and circuit breaker to do their job first, and fallback to intervene only when all have failed. The code calling the pipeline no longer needs any try/catch: it receives either recommendations or an empty list.

Strategy 2: last known good value

For important functionalities, an empty list is not enough. The exchange rate is a perfect example: if the BNR service does not respond, the application cannot show a nonexistent rate but can show the last successfully obtained rate, clearly marked as such.

The pattern is called last known good and has two parts: you save every successful response, and in fallback you serve the saved one.

public sealed record ExchangeRate(
    string Currency, decimal Value, DateOnly PublicationDate, bool IsStale = false);

public sealed class LastKnownGoodStore(IDistributedCache cache)
{
    private static readonly DistributedCacheEntryOptions Retention = new()
    {
        AbsoluteExpirationRelativeToNow = TimeSpan.FromDays(7)
    };

    public Task SetAsync<T>(string key, T value, CancellationToken ct) =>
        cache.SetStringAsync($"lkg:{key}", JsonSerializer.Serialize(value), Retention, ct);

    public async Task<T?> GetAsync<T>(string key, CancellationToken ct)
    {
        var json = await cache.GetStringAsync($"lkg:{key}", ct);
        return json is null ? default : JsonSerializer.Deserialize<T>(json);
    }
}
builder.Services.AddResiliencePipeline<string, ExchangeRate>("bnr-rate", (pipeline, context) =>
{
    var store = context.ServiceProvider.GetRequiredService<LastKnownGoodStore>();

    pipeline
        .AddFallback(new FallbackStrategyOptions<ExchangeRate>
        {
            ShouldHandle = new PredicateBuilder<ExchangeRate>()
                .Handle<HttpRequestException>()
                .Handle<TimeoutRejectedException>()
                .Handle<BrokenCircuitException>(),
            FallbackAction = async args =>
            {
                var last = await store.GetAsync<ExchangeRate>("eur-rate", args.Context.CancellationToken);

                return last is not null
                    ? Outcome.FromResult(last with { IsStale = true })
                    : Outcome.FromException<ExchangeRate>(new RateUnavailableException());
            }
        })
        .AddRetry(retryOptions)
        .AddCircuitBreaker(breakerOptions)
        .AddTimeout(TimeSpan.FromSeconds(3));
});

And in the service consuming the pipeline, each fresh response updates the saved value:

var rate = await _pipeline.ExecuteAsync(
    async token => await _bnrClient.GetEurRateAsync(token), ct);

if (!rate.IsStale)
{
    await _store.SetAsync("eur-rate", rate, ct);
}

return rate;

Three details make the difference between a good fallback and a dangerous one:

  • Explicitly mark old data. The IsStale and PublicationDate fields reach the interface: “BNR rate from 10/08/2026 — today’s rate is not yet available.” The user decides knowingly.
  • Set an age limit. A rate from yesterday is acceptable for display. One from a month ago probably isn’t. The 7-day retention in the example is a business decision, not a technical one.
  • Fail when you have nothing good. If no saved value exists, the fallback throws a clear exception. It does not invent a rate.

A wrong fallback is worse than an error

The last point deserves emphasis. It’s tempting to write a fallback that returns something: a rate of 1, a price of 0, an empty product list in a context where an empty list has meaning. An invoice issued with a rate of 1 leu per euro is not controlled degradation. It’s a wrong invoice, with fiscal consequences, generated automatically and silently.

The rule: a fallback can provide older or fewer data, marked as such. It cannot provide false data. When the only alternative to error is an invented value, error is the correct choice.

Strategy 3: asynchronous decoupling

Back to the two invoicing applications from the beginning. The first was not badly written. It probably had retry, timeout, and maybe circuit breaker. Its problem was architectural: issuing the invoice and transmitting it to ANAF were a single synchronous operation. When ANAF was down, no resilience strategy could save issuing because issuing depended on ANAF.

The second application separated them. Issuing is a local operation: validation, numbering, saving in its own database. Transmission is a separate, asynchronous operation with its own lifecycle. For the user, the invoice goes through visible states:

  1. Issued — saved locally, with allocated number. The synchronous operation stops here.
  2. Waiting to transmit — a message in Outbox, written in the same transaction as the invoice.
  3. Transmitted — uploaded to SPV, with received upload index.
  4. Validated or Rejected — after subsequent status verification.

Transmission is done by a background worker that uses all mechanisms in series:

public sealed class EFacturaUploadWorker(
    IServiceScopeFactory scopeFactory,
    [FromKeyedServices("anaf")] CircuitBreakerStateProvider circuit,
    WorkerHeartbeat heartbeat,
    ILogger<EFacturaUploadWorker> logger) : BackgroundService
{
    protected override async Task ExecuteAsync(CancellationToken stoppingToken)
    {
        using var timer = new PeriodicTimer(TimeSpan.FromSeconds(30));

        while (await timer.WaitForNextTickAsync(stoppingToken))
        {
            heartbeat.Beat();  // liveness: the worker is alive even if ANAF is not

            if (circuit.CircuitState is CircuitState.Open or CircuitState.Isolated)
            {
                continue;  // no point in trying, the circuit is open
            }

            await using var scope = scopeFactory.CreateAsyncScope();
            var uploader = scope.ServiceProvider.GetRequiredService<EFacturaUploader>();

            try
            {
                await uploader.UploadPendingBatchAsync(stoppingToken);
            }
            catch (BrokenCircuitException)
            {
                logger.LogInformation("ANAF unavailable, will retry next cycle.");
            }
        }
    }
}

The CircuitBreakerStateProvider is the same object you gave to the circuit breaker in the anaf pipeline, registered in the container as a keyed service: builder.Services.AddKeyedSingleton("anaf", anafCircuit).

You recognize the parts: the heartbeat from the health checks article, which maintains correct liveness while ANAF is down, the circuit breaker state from article 3, which stops unnecessary calls, and the Outbox from the retry article, which guarantees no invoice is lost.

Decoupling does not completely eliminate risk but moves it to a controllable place. Transmission has a legal deadline, and an invoice that stays too long waiting becomes a real problem. So you add an alert: if there are unsent invoices older than, say, 24 hours, someone must find out. Controlled degradation does not mean ignoring the problem. It means managing it without blocking the user.

Strategy 4: hedging for latency

So far, we have dealt with dependencies that do not respond. There is also a subtler case: dependencies that usually respond quickly but sometimes, unpredictably, very slowly. A p50 of 50 ms but a p99 of 3 seconds. Retry does not help because the request does not fail, it just delays.

Hedging solves this: if the first request has not responded within a short interval, you send another one in parallel, without canceling the first. You use the response that arrives first and cancel the other.

pipeline.AddHedging(new HedgingStrategyOptions<HttpResponseMessage>
{
    MaxHedgedAttempts = 1,
    Delay = TimeSpan.FromMilliseconds(300)  // approximately normal p95
});

For HttpClient there is also AddStandardHedgingHandler(), which can send the parallel request to a different endpoint, for example another region. This also provides tolerance to the failure of an entire region.

Hedging has two strict conditions. First: the operation must be idempotent because you will execute it, at least partially, twice. Practically, hedging is for reads, not writes. Second: it costs extra traffic. With Delay set at p95, about 5% of requests generate a second request. If you set it too low, you double the dependency’s load.

Strategy 5: feature flags as emergency switches

All the above strategies are automatic. Sometimes you also need a manual button: a functionality that consumes too many resources during an incident, an integration with a provider known to have problems, a heavy report you want to stop on Black Friday.

builder.Services.AddFeatureManagement();
{
  "FeatureManagement": {
    "Recommendations": true,
    "AdvancedPdfExport": true,
    "LiveStatisticsDashboard": true
  }
}
public async Task<CheckoutPage> GetCheckoutAsync(int cartId, CancellationToken ct)
{
    var page = await _checkout.BuildAsync(cartId, ct);

    if (await _featureManager.IsEnabledAsync("Recommendations"))
    {
        page.Recommendations = await _recommendations.GetAsync(cartId, ct);
    }

    return page;
}

Combined with a dynamic configuration source, such as Azure App Configuration with automatic refresh, you can disable a feature in seconds without deploying. These flags are called kill switches. The difference from circuit breaker is that there the algorithm decides, and here a person decides, who knows something the algorithm does not: that the provider announced a problem, that a traffic peak is coming, or that a feature must be sacrificed so the rest survive.

Communicate degradation, don’t hide it

The difference between the two invoicing applications was not only technical. The second told the user what was happening. A degraded system that behaves like a healthy one is sometimes more dangerous than one that shows an error.

  • In the interface: status badges on entities (“waiting to transmit”), global banners for affected functionalities, data marked with the last update time.
  • In the API: explicit fields in the response (isStale, asOf), so API clients can make informed decisions.
  • For critical functionalities that really cannot work: 503 Service Unavailable with Retry-After and a clear ProblemDetails, not a generic 500.
  • For the team: every executed fallback is logged and counted. OnFallback in Polly is the right place. A silent fallback running for three days without anyone knowing is a hidden incident, not controlled degradation.
OnFallback = args =>
{
    logger.LogWarning(args.Outcome.Exception,
        "Fallback activated for BNR rate: serving last known value.");
    return default;
}

Summary

  • Graceful degradation is a product decision before it is a technical one. Classify functionalities as critical, important, and optional.
  • Fallback is outside the pipeline: it intervenes only after retry and circuit breaker have failed.
  • Last known good: serves older data, explicitly marked, with an age limit.
  • A fallback can provide older or fewer data, never false. When the only alternative is an invented value, error is the correct choice.
  • Asynchronous decoupling (Outbox + worker) transforms a blocking dependency into a manageable delay.
  • Hedging for unpredictable latency, only on idempotent operations.
  • Feature flags as manual switches for incidents.
  • Communicate degradation in the interface, API, and monitoring.

Resilience in .NET series

  1. Fundamentals and Polly v8
  2. Intelligent retry policies
  3. Circuit breaker and bulkhead
  4. Advanced health checks
  5. Graceful degradation — this article
  6. Chaos engineering basics

In five articles, we built a whole defense system: timeouts, retries, circuit breakers, bulkheads, health checks, and fallbacks. One uncomfortable question remains: how do you know it works? The answer, in the last article of the series: chaos engineering, with error injection strategies integrated into Polly v8 and tests that prove controlled degradation really happens when it should.