↑↓ to navigate Enter to open "…" all these words ANDOR to combine

Run & extract · No. 33

Retries are not a recovery plan: resilience handlers, RTO/RPO, and a restore you actually drilled

The last article promised the resilience handler that makes retries happen, and the recovery objectives behind it. Here is both halves: the Polly pipeline on every outbound client, and the part most "resilience" write-ups skip entirely, a declared RTO/RPO per failure scenario and a restore drill that someone actually ran.


The extraction article ended on a promise: the retry pipeline that keeps those cross-service calls self-healing, plus the recovery objectives behind it. This is that article, and the second half of the promise, recovery, is the one that matters most. (Idempotency, a few articles back, was the other half of the same coin: a retry is only safe if the operation it repeats is idempotent.)

Here is the uncomfortable framing. Almost every "we made it resilient" story stops at the retry policy. Add Polly, add a circuit breaker, set a timeout, ship it. That is the easy 80%. It is also the half that gives you false confidence, because a retry policy answers exactly one question: "what happens when a single call transiently fails?" It says nothing about the questions that actually wake you at 3am: how much data can we lose, how long until we are back, and have we ever proven we can restore at all. A retry policy without a drilled restore is theater. It looks like business continuity from the outside and collapses the first time you need it.

ADR-009 (the row in the primer ADR table reads "Standard resilience handler on every outbound client; declared RTO/RPO + drilled restore") is deliberately two decisions bolted together, because resilience and recovery are two halves of the same posture. This article walks both.

Half one: the standard resilience handler on every outbound client

The runtime half is the part already shown elsewhere in this series, so I will summarize and attribute rather than re-derive it. Every outbound HTTP and gRPC client in the system gets the same Polly pipeline: retry, timeout, and a circuit breaker. It is not wired call-site by call-site. It is configured once, in the Aspire service-side baseline, and every client inherits it.

The wiring lives in AddServiceDefaults() from MMCA.Common.Aspire, the single method every service host calls first in Program.cs. The Aspire one-command deep-dive (Article 31) walks it in full; the short version is that AddServiceDefaults() configures OpenTelemetry, service discovery, and Polly resilience handlers on every HttpClient, with a 30 second per-attempt timeout, a 60 second circuit-breaker sampling window, and a 90 second total-request timeout. Those are the values that article verified against source. Every outbound call rides this pipeline by default. That is the runtime side of ADR-009 stated plainly: a developer adding a new typed client does not opt into resilience and cannot forget it, because it is applied at the HttpClient defaults layer, not per registration.

The gRPC side gets the identical treatment, and the gRPC contracts chapter is explicit about it. AddTypedGrpcClient<TClient>(serviceName) calls AddStandardResilienceHandler(), which gives every gRPC client an explicit Polly pipeline sourced entirely from GrpcResilienceDefaults: the attempt timeout, the total timeout and the retry budget are the same values the HTTP defaults use, and the circuit-breaker values are spelled out on their own (a 0.5 failure ratio, a minimum throughput of 10, a 10 second break), precisely because an east-west gRPC call bypasses the Gateway's active health checks. There is one sharp edge worth knowing, covered in the module-extraction deep-dive (Article 32): a deliberate SocketsHttpHandler override forces explicit HTTP/2 first, because the global resilience handler can otherwise wrap the primary handler in a way that defeats the h2c negotiation. Resilience is then layered back on top. So the order is "force the transport, then add resilience," not the reverse.

This pipeline is not decorative. It is load-bearing for a specific, documented case. In ADC, Conference and Engagement form a bidirectional gRPC pair, and the AppHost gives Engagement a WaitFor(conference) but deliberately gives the reverse Conference-to-Engagement edge only a WithReference with no WaitFor, because a reciprocal wait would deadlock startup (each service waiting for the other to be healthy). The transient "peer not ready" errors that result during startup self-heal through the resilience pipeline. The gRPC chapter says this directly: "the retry plus circuit breaker is what makes them tolerable." That is the runtime payoff. Retries are what let a mutual synchronous dependency boot at all.

Fitness-enforced, not just present

A convention that is "applied everywhere by everyone remembering to" is the same failure mode the idempotency and caching articles kept warning about. The framing across this series is that the standard resilience handler is fitness-enforced: the architecture asserts that outbound clients carry it, rather than trusting each registration. The honest detail is that the primary mechanism for getting it onto every client is structural: it is added at the HttpClient defaults layer in AddServiceDefaults() and inside AddTypedGrpcClient<TClient>, so the only way to get an outbound client without resilience is to bypass the framework's registration helpers entirely. That structural default is what makes "every outbound client" a statement of fact instead of a hope.

The one place resilience left the outbound client: the outbox broker publish

For most of this series, "resilience policy" has meant "outbound HTTP or gRPC client." ADR-087 (accepted 2026-08-18) amends ADR-009 with the first exception, and it is a narrow one: the outbox's broker publish carries a Polly circuit breaker of its own, and nothing else in the persistence path does.

The scope is deliberate. OutboxProcessor holds a per-instance ResiliencePipeline (OutboxProcessor.cs:107, built at :790-801) and wraps exactly one call in it, state.Bus.PublishAsync(state.Event, ct) (:631-635). The in-process dispatch branch sits outside it, because a direct method call into the same process has no transport that can be dead (:627-630, with the else branch at :637-640). The parameters live in BrokerResilienceDefaults: a 0.5 failure ratio (:32), a minimum throughput of 10 attempts before the ratio is evaluated at all (:40), a 30 second sampling window (:47), and a 15 second break (:55). No retry strategy is paired with the breaker (:17-22), because the outbox already is the retry: a Polly retry inside a loop that re-leases and retries would multiply the attempt count without changing the outcome.

Note what this breaker does and does not buy, because it is the cleanest small example of the honesty this article is asking for. It adds no delivery guarantee. BrokenCircuitException is caught by the same catch (Exception ex) as any publish failure (:666), increments RetryCount (:668), re-leases the row with the usual backoff (:677-678), and dead-letters at MaxRetries like everything else (:702). Only the observability differs: a broker.circuit.open.count increment (:691-693) and one log line per batch instead of one per message (:696-700). What it fixes is cost and noise. When the broker is unreachable rather than slow, every message in every batch fails identically, and the breaker converts a batch of connection timeouts into a batch of microsecond rejections. The pipeline is also a per-instance field, so N replicas mean N independent circuits and an effective threshold N times the configured one.

The other half of that record is what it declined to do, and that is the part worth copying. The symmetric move, a breaker on the query path so a failing database sheds load instead of queueing on it, is explicitly rejected. The reason is composition, not appetite: EF Core's EnableRetryOnFailure execution strategy (SQLServerDbContext.cs:63-66, 5 retries, a 10 second maximum delay) already owns retrying at the EF layer, and it constrains how a user-initiated transaction may be written. Wrapping a breaker around a call that is already being retried inside that strategy would either count one logical failure many times or force the strategy to be replaced, which is an EF execution-strategy rework rather than a breaker. So the database keeps EnableRetryOnFailure plus CommandTimeoutSeconds (:55) as its posture, and the asymmetry is recorded as a decision. The ADR's own reason for writing the rejection down is the reusable lesson: an engineer who finds a breaker on the broker and none on the database will otherwise conclude the second was forgotten, and add it.

Half two: declared RTO and RPO per scenario

Now the part that is actually missing from most "resilience" treatments, and the part this article exists to document.

ADR-009 requires each consuming app to declare, in writing, its Recovery Time Objective and Recovery Point Objective per failure scenario. Not "we have backups." Specific numbers, per scenario, that someone signed off on. ADC's infra/DISASTER-RECOVERY.md is the worked example, and it targets three scenarios with explicit figures:

Scenario RPO RTO
Accidental data loss / bad migration (within the PITR window) about 10 minutes up to 2 hours
Single service DB corruption about 10 minutes up to 1 hour (PITR restore-as-new, swap the name)
Full region loss up to 1 hour (geo-redundant replication lag) up to 4 hours (geo-restore plus redeploy)

These numbers exist in source. They are real and modest, and the runbook says why they are modest on purpose: ADC is a regional, non-24x7 conference app, so sub-hour multi-region failover is explicitly declared a non-goal. That sentence is the whole point of declaring objectives. An RTO is not a brag about how fast you are. It is a statement of how fast you have decided to be, so that the architecture, the backup tier, and the on-call expectations can all be sized to match it. A declared four-hour RTO for a region loss is an honest, scoped commitment. An undeclared one is a fight at 3am about what "back up" even means.

The runbook pairs the objectives with the SPOFs it knowingly accepts: one Azure SQL server, one Container Apps environment with every app at minReplicas: 1, one Service Bus namespace, one ACR. Each is listed with its mitigation. Accepting a single point of failure in writing, with the mitigation named, is a different act than having one by accident.

The drilled restore: backups you tested, not backups you assumed

The second word in "declared RTO/RPO and a drilled restore" is the one that separates a posture from a wish.

The backup posture in ADC's runbook is two tiers. Point-in-time restore (PITR) on geo-redundant storage covers the "undo the last bad change" case with an RPO of minutes. Long-term retention (LTR) on all four live per-service databases keeps weekly, monthly, and yearly backups, declared as a serviceDatabaseLtr resource in infra/main.bicep. Those four are the only databases on the SQL server, so the two tiers cover the entire application data estate; the AtlDevCon archive of the legacy shared database lives in blob storage as a bacpac instead, which puts it outside PITR and LTR by design. Backups being declared in infrastructure-as-code, not clicked into existence in a portal, is itself part of the discipline: the retention policy is version-controlled and reviewable.

But a backup you have never restored is a hypothesis, not a recovery plan. So the runbook ships an Azure-CLI drill script (az sql db restore, then az sql db show, then az sql db delete): it PITR-restores a throwaway copy of a database, times the restore as the measured RTO, verifies the copy comes back Online, then deletes it. The automated check is deliberately ARM-level (is the restored database Online?); a deeper data check, row counts against the original via the SQL admin login, is a documented manual step you run with -KeepCopy. The objective is not to recover anything. It is to prove the restore path works before you need it under pressure, and to time it against the declared RTO. ADR-009 requires the drill to be recorded in a drill-result table.

The drill is executed, repeated and enforced. Its first recorded run, on 2026-06-20, took ADC_Conference through a point-in-time restore to a throwaway copy, verified and timed at 2.6 minutes against the declared 2-hour RTO, recorded in the drill table in DISASTER-RECOVERY.md and closing TD-10. It runs as a dr-drill.yml workflow over a dr-restore-drill.ps1 script, on a weekly cron (Mondays 06:00 UTC) as well as on-demand workflow_dispatch, and the deploy pipeline carries a dr-freshness job in its needs that reads the latest successful drill and fails the deploy if it is older than an 8-day window, so a stale recovery proof blocks the deploy instead of being a soft reminder. The ADC scorecard puts resilience at its top maturity precisely on that gate. Coverage is a rotation rather than a single-database story: the scheduled run picks its target across all four live per-service databases by ISO week number, so each one earns its own recovery proof roughly monthly, and the drill table records six passing restores from 2026-06-20 through 2026-08-10 covering ADC_Identity, ADC_Conference, ADC_Engagement, and ADC_Notification, the slowest at 4.4 minutes against the 2-hour RTO. That is also the cleanest illustration of why "drilled" is a separate requirement from "documented": a runbook can be complete and correct and still be a hypothesis, and only a run that is timed and written down turns a correct procedure into a measured recovery posture. What remains genuinely open is narrower: every one of those proofs is the same scenario, a single-database point-in-time restore, so the region-loss objective and the full multi-service restore order are declared targets rather than measured ones, and the drill plus its freshness gate live in the consuming app's pipeline, because ADR-009 makes recovery each app's obligation rather than something the framework can run for it.

Database-per-service multiplies the restore surface

There is a structural reason recovery is harder here than in a monolith, and it is the direct cost of a decision celebrated elsewhere in this series.

ADR-006 (database-per-service) gives every service its own database and its own outbox, the right call for isolation and the outbox-race fix. But it has a recovery consequence: there is no single database to restore. There are four (ADC_Identity, ADC_Conference, ADC_Engagement, ADC_Notification), each with its own PITR window, its own LTR policy, and its own place in the restore order. A point-in-time restore of one service to 10:42 does not automatically align with the others, and cross-service references are scalar IDs with consistency flowing through the outbox and broker, so a restore can leave an event mid-flight. The runbook's per-database, restore-to-new-name-then-repoint procedure is the operational answer, and it is more steps than "restore the database." Database-per-service multiplies the backup and restore surface. That is not a reason to avoid it. It is a reason to declare the objectives per service and to drill the restore, which is precisely what ADR-009 demands.

The DR runbook and the deploy-time auto-revert

The runbook governs four per-service databases, ADC_Identity, ADC_Conference, ADC_Engagement and ADC_Notification, and those four are the entire application data estate: one Basic database per service on the same SQL server, each owned and migrated by exactly one service. The AtlDevCon archive of the legacy shared database is not a database at all. It is a single bacpac blob, sql-archive/AtlDevCon-20260902.bacpac, held in the deployment's storage account and restorable onto that SQL server with az sql db import in about ten minutes, with its own runbook carrying the exact commands. That blob is the last-resort source of record, and holding it as a blob is what takes it out of the PITR and LTR scope: a static archive nothing writes to does not need a retention policy, it needs a documented import path and a storage account that still accepts the authorization mode that import command uses.

The deploy pipeline closes the loop on the runtime side with two post-deploy gates that answer different questions. The activation gate asserts that for every app the newest revision reports healthy, running, and carrying 100% of the traffic, which is what proves the code just built is the code currently serving. The reachability gate then probes every service through the Gateway: /health, the JWKS endpoint, an anonymous GET /Events on Conference, the auth-gated /Bookmarks and /Notifications/inbox (where a 401 is the healthy signal, because it proves the request traversed Gateway to service to auth pipeline), and the UI root. Either gate failing rolls every app back to its last healthy provisioned revision and fails the job. The order matters: HTTP probes all enter through the Gateway, and a healthy Gateway keeps serving from the previous backend revision when the new one never goes ready, so every probe answers from old code and the run goes green, and the activation gate is what catches that. Together the two gates are recovery automation for the most common failure (a bad deploy), beside the manual restore procedures for the rarer, scarier ones.

Trade-offs, honestly

  • RTO/RPO are only as real as the last drill. ADC's objectives are declared with specific numbers, and the restore is drilled repeatedly: the drill table records six passing restores from 2026-06-20 through 2026-08-10, covering every live database on a weekly rotation, from a 1.8-minute ADC_Identity restore to a 4.4-minute one, all well inside the 2-hour RTO. That is six timed proofs of a single scenario: a one-database point-in-time restore. The cadence is enforced rather than aspirational: the drill runs on a weekly cron, and a dr-freshness deploy gate blocks the deploy when the latest successful drill goes stale. The recovery time is measured rather than assumed for that path, but the figures for the other scenarios (region loss, the full multi-service restore order) remain declared targets until each is drilled too.
  • Database-per-service multiplies the backup and restore surface. Four databases mean four PITR windows, four LTR policies, and a restore order to get right. The isolation win from ADR-006 is paid for in recovery complexity. The mitigation is per-service objectives and a per-database restore procedure, not a single "restore everything" button, because that button cannot exist across independent databases.
  • Retries can amplify load. The resilience handler resends calls that did not visibly succeed, which is the whole reason the idempotency article exists: a retry is only safe if the operation it retries is idempotent. Under a partial outage, aggressive retries across many clients can also pile load onto a struggling dependency. The circuit breaker (the 60 second sampling window) is what bounds that, but the tension is real and the timeouts are tuned, not arbitrary.
  • Single-region by design. ADC accepts one SQL server, one Container Apps environment, one Service Bus namespace, and one ACR as written-down single points of failure. The four-hour region-loss RTO is honest about this. A system that needs sub-hour multi-region failover would declare different objectives and pay for them; this one declared that it does not, and sized accordingly.

Apply this even without MMCA

The posture ports to any stack, and most of it is discipline rather than code:

  1. Put the resilience handler in one shared place every client inherits (the HttpClient/gRPC defaults layer), so no caller can ship without it. Make "every outbound client has it" structurally true, not a code-review checklist item.
  2. Declare RTO and RPO per failure scenario, in writing, with numbers. "We have backups" is not an objective. "Up to 2 hours, RPO 10 minutes, for a bad migration" is.
  3. Write down the single points of failure you are accepting, each with its mitigation. An accepted SPOF is an engineering decision; an unnoticed one is an incident.
  4. Drill the restore and record the result. A backup you have never restored is a hypothesis. The drill is what converts it into a recovery plan, and the recorded result is what converts it into a timed one.
  5. Account for your data topology. If you split databases per service, your restore surface multiplied; declare and drill per service, not once for the whole system.

The takeaway: resilience is not retries; it is a declared, tested recovery posture. A Polly pipeline keeps a single call alive under transient failure. An RTO, an RPO, and a restore you actually drilled keep the business alive after a real one. Ship both, and write down the drill.


What we covered: how ADR-009 bolts two decisions together; the standard Polly resilience handler (retry, timeout, circuit breaker) applied to every outbound HTTP and gRPC client through AddServiceDefaults() and AddTypedGrpcClient, and why it is what lets a bidirectional gRPC pair even boot; the one place ADR-087 extends that posture past outbound clients, a circuit breaker around the outbox broker publish and nothing else, plus the per-query database breaker the same record rejects because it does not compose with EF's execution strategy; why retries alone are insufficient; the declared RTO/RPO per scenario in ADC's DR runbook; the two-tier PITR/LTR backup posture and the drills that were run and timed rather than assumed (a 2.6-minute ADC_Conference PITR restore on 2026-06-20 closed TD-10, and a weekly rotation records six proofs through 2026-08-10 covering all four live databases); how database-per-service multiplies the restore surface; and the activation and reachability smoke gates that surround the runbook.

Next in the series: architecture fitness functions, the executable rules that keep every pattern in this series honest by failing the build when one is violated.

MMCA.Common is Apache-2.0 licensed and open source. Star the repo, read the operational-runbooks chapter of the onboarding guide, or dotnet add package MMCA.Common.Aspire and try it.

Tags: .NET, Resilience, Software Architecture, DevOps, Reliability