How retries really work with the Service Bus trigger

ยท 2 min read

When a Service Bus-triggered function throws, the message is retried immediately, ten times in a row, and then dead-lettered. If you need a delay between attempts, you have to build it.

A Service Bus-triggered Azure Function calls a partner API that's down for maintenance. The function throws. What happens next surprises a lot of people.

The default: immediate redelivery

When the function fails, the message is abandoned. Its lock is released, its delivery count goes up, and it's available again straight away. The function picks it up again within milliseconds and fails again. That repeats until the delivery count reaches the queue's MaxDeliveryCount (10 by default), at which point the message moves to the dead-letter queue.

So a five-minute outage at the partner turns into ten failures in a couple of seconds and a dead-lettered message. No backoff at all.

Azure Functions has its own retry policies with fixed or exponential delays, but they don't apply to the Service Bus trigger, which relies on Service Bus's delivery count instead.

Sort failures into three kinds

Permanent failures, such as invalid data or a record that doesn't exist. Retrying won't help. Dead-letter immediately with a reason:

await actions.DeadLetterMessageAsync(message,
    deadLetterReason: "InvalidPayload", deadLetterErrorDescription: ex.Message, cancellationToken: ct);

Brief transient failures, such as a dropped connection or a single timeout. Retry inside the function with a resilience handler on the HttpClient, a few attempts over a few seconds, before the function fails at all.

Longer outages, where you need minutes between attempts. That needs a scheduled retry.

Scheduled retries

Instead of failing, complete the original message and schedule a copy for later, with an attempt counter in its properties:

catch (HttpRequestException) when (Attempt(message) < 5)
{
    var attempt = Attempt(message) + 1;
    var retry = new ServiceBusMessage(message)   // copies body and properties
    {
        ScheduledEnqueueTime = DateTimeOffset.UtcNow.AddMinutes(Math.Pow(2, attempt))   // 2, 4, 8, 16, 32 minutes
    };
    retry.ApplicationProperties["RetryAttempt"] = attempt;

    await retrySender.SendMessageAsync(retry, ct);
    await actions.CompleteMessageAsync(message, ct);
}

static int Attempt(ServiceBusReceivedMessage m) =>
    m.ApplicationProperties.TryGetValue("RetryAttempt", out var value) ? Convert.ToInt32(value) : 0;

retrySender is a ServiceBusSender for the same queue, registered through dependency injection. After the last attempt, the exception propagates as usual and the message eventually dead-letters.

Two cautions. The copy keeps the original's MessageId, so if duplicate detection is turned on for the queue, give it a new one, such as the original ID plus the attempt number, or the broker drops it as a duplicate. And sending the copy and completing the original aren't atomic, so a crash in between can produce a duplicate. Handlers must be idempotent either way.

Tune the host

In host.json, under extensions.serviceBus, two settings matter most:

  • maxConcurrentCalls limits how many messages each instance processes at once. Lower it if the downstream system can't take the load.
  • maxAutoLockRenewalDuration controls how long the host keeps renewing a message's lock. It should comfortably exceed your longest expected processing time, or the message is redelivered while still being processed.

Takeaway

The Service Bus trigger retries failed messages immediately until the delivery count runs out. Dead-letter permanent failures straight away, retry brief glitches in-process, and schedule delayed copies of the message when you need real backoff.