Monday morning, 9:02. The database performs a planned failover, which lasts 20 seconds. Nothing serious: that's what retries are for. But retries exist everywhere. The API Gateway retries 3 times. The order service retries 3 times. The repository in the order service also retries 3 times. Each user request becomes up to 64 calls to a database that doesn't respond anyway.
At 9:02:20 the database comes back. Exactly then it is hit by a synchronized wave of requests, all retried after the same 2-second pause, then 4, then 8. It fails again. The 20-second failover becomes a 40-minute incident.
Retry is the simplest resilience strategy and, for the same reason, the easiest to misuse. A good retry recovers a transient failure without the user noticing anything. A bad retry turns a small problem into an attack on your own infrastructure. In this second article in the series, we see what makes the difference.
The first question: is it worth retrying?
In the first article we classified failures into four categories. Retry makes sense only for one of them: transient failures. The first thing you configure is not the number of retries, but the predicate that decides what is retried.
- Yes:
HttpRequestException(connection refused or reset, DNS),TimeoutRejectedExceptionfrom per-attempt timeout,408 Request Timeout,429 Too Many Requests,502 Bad Gateway,503 Service Unavailable,504 Gateway Timeout. - Debatable:
500 Internal Server Error. Sometimes it's transient (a deadlock behind the API), other times it's a bug that will fail identically every time. The standard handler inMicrosoft.Extensions.Http.Resilienceretries it. For APIs you know well, you can decide otherwise. - No:
400,401,403,404,409,422. The request is wrong, not the network. Retrying only delays the error and consumes resources.
In Polly v8, the predicate is written with PredicateBuilder. For HTTP calls you use the generic variant, which can inspect the result as well as exceptions:
static bool IsTransient(HttpResponseMessage response) =>
response.StatusCode is HttpStatusCode.RequestTimeout
or HttpStatusCode.TooManyRequests
or HttpStatusCode.InternalServerError
or HttpStatusCode.BadGateway
or HttpStatusCode.ServiceUnavailable
or HttpStatusCode.GatewayTimeout;
var shouldHandle = new PredicateBuilder<HttpResponseMessage>()
.Handle<HttpRequestException>()
.Handle<TimeoutRejectedException>()
.HandleResult(IsTransient);
Backoff: how long you wait between attempts
An immediate retry after failure is almost always useless. If the service responded with 503 three milliseconds ago, the chances it is healthy now are minimal. You need a pause, and Polly v8 offers three types via DelayBackoffType:
- Constant — the same pause every time: 1s, 1s, 1s. Useful for polling or very short local operations.
- Linear — the pause increases by a fixed step: 1s, 2s, 3s.
- Exponential — the pause doubles: 1s, 2s, 4s, 8s. The suitable variant for almost any external dependency, because it gives the backend service more and more time to recover as the problem seems more serious.
Exponential backoff has an obvious problem: it grows very fast. The tenth retry with a base of 1 second would wait over 8 minutes. That's why there is MaxDelay, which caps any pause regardless of the formula.
Jitter: why it is not optional
Return to the scenario at the beginning. A thousand client instances receive an error in the same second because the dependency went down for everyone at once. All have the same exponential backoff without jitter. What happens?
They all retry exactly after 1 second. Then all exactly after another 2. Then all exactly after another 4. The load does not spread over time, but comes in synchronized waves, each with the full force of the thousand clients. The service trying to recover receives the biggest hit exactly when it is most fragile.
Jitter adds a random component to each pause, so that client retries spread out over time instead of overlapping. The wave of 1,000 requests in the same millisecond becomes a flow of a few hundred requests per second, which a recovering service can absorb.
In Polly v8, jitter is a single property. For exponential backoff, Polly uses a decorrelated jitter algorithm, which keeps the exponential growth on average but distributes pauses uniformly without clustering:
var pipeline = new ResiliencePipelineBuilder<HttpResponseMessage>()
.AddRetry(new RetryStrategyOptions<HttpResponseMessage>
{
ShouldHandle = shouldHandle,
MaxRetryAttempts = 3,
BackoffType = DelayBackoffType.Exponential,
Delay = TimeSpan.FromMilliseconds(500),
MaxDelay = TimeSpan.FromSeconds(10),
UseJitter = true
})
.Build();
The rule is simple: if a retry can be executed simultaneously by multiple clients, and it almost always can, UseJitter = true. There is almost no production scenario where you want synchronized retries.
Retry-After: when the service tells you how long to wait
Some services explicitly tell you when to come back. A 429 from a rate-limited API or a 503 during maintenance often come with the Retry-After header, either as a number of seconds or as a date. Ignoring it and retrying after your own backoff means hitting again a service that clearly told you it cannot serve you yet. For some public APIs, this can also lead to temporary blocking of the access key.
The standard handler in Microsoft.Extensions.Http.Resilience already respects this header (ShouldRetryAfterHeader is enabled by default). When building the pipeline manually, you use DelayGenerator:
DelayGenerator = args =>
{
if (args.Outcome.Result?.Headers.RetryAfter is { } retryAfter)
{
TimeSpan? delay = retryAfter.Delta
?? (retryAfter.Date is { } date ? date - DateTimeOffset.UtcNow : null);
if (delay > TimeSpan.Zero)
{
return ValueTask.FromResult(delay);
}
}
// null = use configured backoff
return ValueTask.FromResult<TimeSpan?>(null);
}
One thing to check: if the service asks you to wait 60 seconds, but the total operation timeout is 10, retrying makes no sense. In such cases, it is healthier to fail immediately and let a higher layer, for example a message queue, resume the operation later.
Idempotency: the question you must ask before any retry
Suppose you send an invoice to an external API. The request goes out, the server processes it, saves the invoice, but the response is lost on the way due to a connection reset. For you, it's a HttpRequestException, a classic transient failure. You retry. The server receives the second request and creates a second invoice.
The problem is not the retry itself. The problem is that you retried an operation that is not idempotent. An operation is idempotent if executing it multiple times has the same effect as executing it once.
- By HTTP semantics,
GET,HEAD,PUT,DELETE, andOPTIONSare idempotent. This assumes the API respects the semantics, which is worth verifying, not assuming. POSTandPATCHare not idempotent by default.
For clients that send unsafe operations, the standard handler allows disabling retry on those methods:
builder.Services
.AddHttpClient<IInvoiceClient, InvoiceClient>()
.AddStandardResilienceHandler(options =>
{
options.Retry.DisableForUnsafeHttpMethods(); // POST, PATCH, PUT, DELETE, CONNECT
});
A better solution, when you control the server as well, is to make the operation idempotent.
Idempotency-Key
The model used by Stripe, many banking APIs, and formalized in an IETF draft is simple: the client generates a unique key for the business operation and sends it in a header. The server remembers processed keys and, on a repeated request, returns the saved response instead of executing the operation again.
var request = new HttpRequestMessage(HttpMethod.Post, "api/invoices")
{
Content = JsonContent.Create(invoice)
};
// The key identifies the operation, not the attempt: it is generated ONCE
request.Headers.Add("Idempotency-Key", invoice.Id.ToString());
var response = await httpClient.SendAsync(request, cancellationToken);
The critical detail is in the comment. If you generate the key inside the retry logic, each attempt has a new key and protection disappears. The safest is to derive the key from an identifier that already exists, such as the invoice ID generated at creation. The resilience handler resends the same HttpRequestMessage, so the header is preserved on each attempt.
On the server you need a table where the key has a uniqueness constraint:
CREATE TABLE IdempotencyKeys
(
TenantId UNIQUEIDENTIFIER NOT NULL,
IdempotencyKey NVARCHAR(100) NOT NULL,
RequestHash VARBINARY(32) NOT NULL,
StatusCode INT NULL, -- NULL = processing
ResponseBody NVARCHAR(MAX) NULL,
CreatedAt DATETIME2 NOT NULL,
CONSTRAINT PK_IdempotencyKeys PRIMARY KEY (TenantId, IdempotencyKey)
);
The flow on receiving a request:
- You try to insert the key with
StatusCode = NULL. The uniqueness constraint makes the step atomic, even if two identical requests arrive in the same millisecond. - If the insert succeeds, you are first. You execute the operation, then save the status code and response.
- If the insert fails on duplicate key and there is a saved response, you return it exactly as sent the first time.
- If the key exists but
StatusCodeis stillNULL, the original operation is still in progress. You return409 Conflict, and the client will retry later. - If the key exists but
RequestHashdiffers, the client reused the key for another request. You return422: it's a client bug, not a retry.
Keys can be deleted after 24–48 hours. No legitimate retry lasts that long.
Retry storms: when resilience becomes an attack
Let's return to the 64 calls from the beginning. The math is brutal: if each layer does one attempt plus 3 retries, each layer multiplies traffic by 4. Three layers means 4 × 4 × 4 = 64. Exactly when the final dependency most needs quiet, it receives 64 times more traffic than normal.
Three rules to prevent this:
- Retry in only one layer. Usually, the layer closest to the failing dependency. Upper layers propagate the error quickly without retrying.
- Don't put retries over SDKs that already retry. EF Core with
EnableRetryOnFailure, Cosmos DB SDK (which automatically retries 429 requests), Azure SDKs: all have their own policies. A Polly pipeline over them multiplies attempts without you knowing. - Use a retry budget. Limit retries to a percentage of total traffic, not just a number per request.
Retry budget: a global cap
MaxRetryAttempts = 3 limits retries per request. It does not limit retries in total. When 100% of requests fail, each will do all 3 retries, and traffic increases 4 times.
A retry budget solves this. The idea comes from gRPC's retry throttling mechanism: a token bucket that decreases on each failure and slowly increases on each success. Retries are allowed only as long as the bucket is at least half full. When the dependency is completely down, the bucket empties quickly and retries stop by themselves. When it recovers, successes refill it.
public sealed class RetryBudget(double maxTokens = 10, double tokenRatio = 0.1)
{
private readonly Lock _lock = new();
private double _tokens = maxTokens;
public void RecordSuccess()
{
lock (_lock) _tokens = Math.Min(maxTokens, _tokens + tokenRatio);
}
public bool TryRecordFailureAndAllowRetry()
{
lock (_lock)
{
_tokens = Math.Max(0, _tokens - 1);
return _tokens > maxTokens / 2;
}
}
}
Integration into Polly is done directly in the ShouldHandle predicate, which is evaluated after each attempt. The budget is registered as a singleton to be shared by all requests to the same dependency:
builder.Services.AddSingleton<RetryBudget>();
builder.Services.AddResiliencePipeline<string, HttpResponseMessage>("provider", (pipeline, context) =>
{
var budget = context.ServiceProvider.GetRequiredService<RetryBudget>();
pipeline.AddRetry(new RetryStrategyOptions<HttpResponseMessage>
{
MaxRetryAttempts = 3,
BackoffType = DelayBackoffType.Exponential,
UseJitter = true,
ShouldHandle = args =>
{
bool transient = args.Outcome switch
{
{ Exception: HttpRequestException or TimeoutRejectedException } => true,
{ Result: { } response } => IsTransient(response),
_ => false
};
if (!transient)
{
budget.RecordSuccess();
return PredicateResult.False();
}
return ValueTask.FromResult(budget.TryRecordFailureAndAllowRetry());
}
});
});
With default values (10 tokens, 0.1 added for each success), retries stop after about 5 consecutive failures and resume after about 10 successful requests per lost token. Practically, retries cannot exceed about 10% of successful traffic. The example above treats any non-transient response as success; you can be stricter if you want.
In the next article we look at a related but much more radical mechanism: the circuit breaker, which not only stops retries but completely stops calls to a sick dependency.
Retry in EF Core: execution strategy
For SQL Server, EF Core has its own retry mechanism, which knows specific transient error codes: deadlocks, Azure SQL failovers, resource limits.
builder.Services.AddDbContext<AppDbContext>(options =>
options.UseSqlServer(connectionString, sql =>
sql.EnableRetryOnFailure(
maxRetryCount: 5,
maxRetryDelay: TimeSpan.FromSeconds(10),
errorNumbersToAdd: null)));
The trap appears when you use explicit transactions. With retry enabled, the following code throws InvalidOperationException, because EF Core cannot retry only part of a transaction you opened:
// WRONG with EnableRetryOnFailure
await using var tx = await db.Database.BeginTransactionAsync(ct);
The correct variant is to put the entire transaction inside the execution strategy so that a failure retries the whole block from the start:
var strategy = db.Database.CreateExecutionStrategy();
await strategy.ExecuteAsync(async () =>
{
await using var tx = await db.Database.BeginTransactionAsync(ct);
db.Invoices.Add(invoice);
db.Outbox.Add(OutboxMessage.From(invoice));
await db.SaveChangesAsync(ct);
await tx.CommitAsync(ct);
});
Because the block can run multiple times, it must be repeatable: no external effects (emails, HTTP calls) inside it. This is exactly one of the reasons why the Outbox pattern in the example works so well.
Time budget: retry and timeout must agree
Returning to the configuration from the first article: per-attempt timeout of 2 seconds, total timeout of 10 seconds, 3 retries with exponential backoff from 500 ms. The worst case:
- 4 attempts × 2 seconds = 8 seconds
- pauses: approximately 0.5 + 1 + 2 = 3.5 seconds (variable due to jitter)
- total: approximately 11.5 seconds, thus over the total timeout of 10 seconds
Not necessarily a mistake, but it must be a conscious decision: in the worst case, the third retry will never happen because the total timeout cuts the operation short. If you want all retries to have a chance to run, increase the total timeout or reduce the number of retries. If the user waits for a response on a page, usually the opposite is better: fewer retries and a fast failure, followed by controlled degradation (the subject of article 5).
Observability: log every retry
A successful retry is invisible to the user, which is good. But if it is also invisible to you, you won't know that a dependency is degrading until retries stop succeeding. OnRetry is the right place for structured logging:
OnRetry = args =>
{
logger.LogWarning(
"Retry {Attempt} after {Delay} ms. Cause: {Cause}",
args.AttemptNumber + 1,
args.RetryDelay.TotalMilliseconds,
args.Outcome.Exception?.GetType().Name
?? ((int?)args.Outcome.Result?.StatusCode)?.ToString());
return default;
}
Combined with Polly metrics in OpenTelemetry, you have a very good indicator of dependency health: the ratio of retries to requests. When it rises from 0.1% to 5%, you have a problem that is not yet visible in errors but will soon be.
Summary
- Retry only transient failures. The
ShouldHandlepredicate is the most important part of configuration. - Exponential backoff, with
MaxDelayand always withUseJitter = true. - Respect
Retry-Afterwhen the service sends it. - Don't retry operations that are not idempotent. Make them idempotent with an
Idempotency-Keygenerated once per operation. - Retry in only one layer and don't double retries of SDKs.
- A retry budget limits retries globally, not just per request.
- In EF Core, explicit transactions run inside
CreateExecutionStrategy().ExecuteAsync(...). - Check if the sum of attempts and pauses fits within the total timeout.
Resilience in .NET Series
- Fundamentals and Polly v8
- Smart retry policies — this article
- Circuit breaker and bulkhead
- Advanced health checks
- Graceful degradation
- Chaos engineering basics
In the next article we look at what to do when retries no longer help: the circuit breaker, with Closed, Open, and Half-Open states, and the bulkhead, which prevents a single sick dependency from consuming all application resources.