Run & extract · No. 32
Extracting a module to a gRPC service, live

A step-by-step walkthrough of lifting a module out of the monolith into its own gRPC service without rewriting application code. Define a contract, wire a typed client, flip the message bus to a broker, federate auth with JWKS, orchestrate with Aspire, route through a gateway. And it stays reversible.
"We will start as a monolith and split into services later" is one of the most common architecture promises, and one of the most commonly broken. The split gets deferred because, when the day comes, "later" turns out to mean a rewrite: the application code is tangled with the transport, the calls between modules are direct method calls that assume a shared process, and pulling one module out means touching everything.
MMCA.Common's thesis is that the split should be a wiring change, not a rewrite, and the way you earn that is by holding one invariant from day one. This article walks through extracting a module live, step by step, and the whole walkthrough only works because of step zero.
Step 0: the starting invariant
The precondition that makes everything else possible: application and domain code already talk to abstractions, and transport choices live at the edges.
Concretely, that means two things were true while the code still ran as a monolith:
- Asynchronous, fire-and-forget flows go through
IMessageBus, defined inMMCA.Common.Application. The application never references the broker library. - Synchronous cross-module calls go through a plain C# interface (for example
IBookmarkCountService), resolved from DI. The caller does not know whether the implementation is an in-process object or something else.
This invariant is not a hope. It is enforced. MicroserviceExtractionTests in the architecture suite
fail the build if MassTransit, gRPC, or Protobuf types appear in any Domain, Application, or Shared
assembly. So you cannot accidentally couple the core to a transport during normal development, which
means when extraction day arrives the core is genuinely host-agnostic. If your codebase does not have
this property, stop here and fix that first. Everything below assumes it.
We will extract a module (call it Engagement) whose interface IBookmarkCountService a peer module
(Conference) calls synchronously.
Step 1: define the .Contracts project
The wire surface lives in its own project whose name ends in .Contracts. That suffix is load-bearing.
Directory.Build.props gives any *.Contracts project special treatment: it auto-pulls Grpc.Tools
and Google.Protobuf and compiles every Protos/**/*.proto with GrpcServices="Both", generating
both the server base class and the client stub. One shared package serves the producer and the
consumer.
MMCA.App.Engagement.Contracts/
├── MMCA.App.Engagement.Contracts.csproj # name ends in .Contracts -> proto auto-compile
└── Protos/
└── bookmark_count.proto # GrpcServices="Both" applied by Directory.Build.props
// Protos/bookmark_count.proto
syntax = "proto3";
service BookmarkCountService {
rpc GetCount (GetCountRequest) returns (GetCountResponse);
}
message GetCountRequest { int32 session_id = 1; }
message GetCountResponse { int32 count = 1; }
The contract project also ships a hand-written gRPC adapter that implements the same C# interface
the modules already used in-process. The adapter holds the generated client and translates an interface
call into a gRPC call and the proto response back into the C# return type. Because both the in-process
implementation and the adapter satisfy the identical IBookmarkCountService, swapping monolith for
microservice is a registration change. The application code that calls IBookmarkCountService never
changes.
Step 2: server side, AddGrpcServiceDefaults()
In the extracted Engagement service's host, register the gRPC server defaults and map the service:
// Engagement.Service Program.cs
builder.Services.AddGrpcServiceDefaults(); // from MMCA.Common.Grpc
...
app.MapGrpcService<BookmarkCountGrpcService>();
AddGrpcServiceDefaults() registers GrpcResultExceptionInterceptor plus gRPC reflection.
The interceptor is the piece that keeps your error model intact across the wire. Your gRPC service
implementation calls the inner C# service, gets back a Result, and calls result.ThrowIfFailure().
That throws a ResultFailureException carrying the Error list; the interceptor catches it for all four
server call shapes, logs it, and rethrows errors.ToRpcException(), an RpcException whose
StatusCode comes from a FrozenDictionary<ErrorType, StatusCode> mapping that mirrors the HTTP
ErrorType-to-status table in ErrorHttpMapping, the internal type ApiControllerBase reads on the
REST side (the unhandled-failure filter reads the same one). The trailers carry every error as
structured metadata. The client adapter reads those trailers back and reconstructs
Result.Failure(errors). One error model, two transports, written once in the interceptor.
Step 3: client side, AddTypedGrpcClient<TClient>(serviceName)
On the calling side (Conference's host), wire the typed gRPC client and replace the local interface registration with the adapter:
// Conference.Service Program.cs, AFTER ModuleLoader has run
services.AddEngagementBookmarkCountClient(); // ships in the .Contracts project; this:
// 1. calls AddTypedGrpcClient<BookmarkCountServiceClient>("engagement")
// 2. services.Replace(IBookmarkCountService -> BookmarkCountServiceGrpcAdapter)
AddTypedGrpcClient<TClient>(serviceName) wires the generated client to Aspire service discovery,
addressing http://{serviceName} over HTTP/2 cleartext (h2c) with prior knowledge. Two interceptor
and resilience details ride along automatically:
- A
JwtForwardingClientInterceptorcopies the inboundAuthorizationheader off the currentHttpContextonto the outgoing call's metadata, so the caller's JWT rides along downstream and distributed authorization works without each handler threading a token (it is a no-op outside an HTTP request, for example in a background processor). AddStandardResilienceHandler()gives every gRPC client the same Polly retry, timeout, and circuit-breaker pipeline as the HTTP clients. A deliberateSocketsHttpHandleroverride forces explicit HTTP/2 so the resilience handler cannot defeat the h2c negotiation.
The services.Replace(...) is deliberate and uses Replace, not TryAdd. By the time this runs (after
ModuleLoader), the container already holds either the real in-process IBookmarkCountService (if
Engagement is enabled in this host) or a Disabled* stub (registered by the module's
RegisterDisabledStubs() when the peer is disabled). Replace wins over both, so after the call the
resolved interface is always the gRPC adapter pointing at the extracted peer. That is why the call comes
after module discovery.
Step 4: flip the message bus from in-process to broker
Synchronous calls now cross the wire. The asynchronous, fire-and-forget flows (domain and integration events) need the same treatment, and the framework makes it a configuration switch rather than a code change.
IMessageBus has two implementations in Infrastructure: InProcessMessageBus (in-monolith, delivery is
a method call) and BrokerMessageBus (a MassTransit broker, RabbitMQ locally and Azure Service Bus in
production, with an IntegrationEventConsumer).
MessageBusSettings selects the mode. In the monolith you ran in-process; in the extracted service you
flip the setting to the broker.
// Program.cs: the broker registration is conditional on settings, not hard-coded.
builder.Services.AddBrokerMessaging(builder.Configuration);
// MessageBus__Provider=RabbitMq (injected by the AppHost's WithBroker) selects BrokerMessageBus.
Your aggregate code, which does AddDomainEvent(new BookmarkAdded(...)), does not change. The
transactional outbox is what makes the switch safe: the same durable outbox row is drained to an
in-process handler today and published to the broker tomorrow. The reliability pattern and the
extraction boundary are the same mechanism (ADR-003).
Step 5: JWKS so the new service validates tokens
The extracted Engagement service now receives forwarded JWTs from Conference and must validate them. It
does so against the issuer's JWKS, the public signing keys, not a shared secret. IJwksProvider
(implemented by RsaJwksProvider in Infrastructure) exposes the signing keys, and
JwksEndpointExtensions in the API layer serves /.well-known/jwks.json so any extracted service can
fetch the issuer's public keys and validate independently. JWKS discovery is routed through the gateway
(ADR-004). No service ever needs to share a symmetric secret with another, which is what makes the fleet
safe to grow.
Step 6: wire it in the AppHost with MMCA.Common.Aspire.Hosting
Local orchestration is where the topology becomes runnable. MMCA.Common.Aspire.Hosting provides
AppHost extension methods for the cross-cutting infrastructure of an extracted deployment: the RabbitMQ
broker, JWKS service discovery, and the per-service data sources. The gRPC edge is not one of them, as
the snippet below shows: peers are wired with stock Aspire WithReference.
// AppHost Program.cs
var broker = builder.AddRabbitMQ("broker");
var engagementService = builder.AddProject<Projects.MMCA_App_Engagement_Service>("engagement")
.WithBroker(broker); // .WithReference(broker).WaitFor(broker) + MessageBus__Provider=RabbitMq
var conferenceService = builder.AddProject<Projects.MMCA_App_Conference_Service>("conference")
.WithBroker(broker);
// gRPC edge: Conference calls Engagement. WithReference injects services__engagement__http__0
// so AddTypedGrpcClient<...>("engagement") can resolve http://engagement at runtime.
conferenceService.WithReference(engagementService);
engagementService.WithReference(conferenceService).WaitFor(conferenceService);
WithBroker chains .WithReference(broker).WaitFor(broker).WithEnvironment("MessageBus__Provider", "RabbitMq"), so each service waits for RabbitMQ to be healthy and gets the env var that makes
AddBrokerMessaging() select the broker transport. WithReference on a project resource injects the
services__{name}__http__0 discovery variables the typed gRPC client uses to resolve http://{name}.
One sharp edge worth internalizing: a bidirectional gRPC pair (Conference calls Engagement and
Engagement calls Conference) must not have reciprocal WaitFor calls, or startup deadlocks (each waits
for the other to be healthy). Declare WaitFor in one direction only; the transient "peer not ready"
errors on the other edge self-heal through the resilience pipeline you got for free in step 3 (ADR-007).
Step 7: route through the YARP gateway
Clients never address a service directly. A single YARP reverse-proxy gateway owns the route-to-service map and is the only entry point. It has no DbContext and no controllers; it is CORS, static files, the route table, and the edge concerns in step 8. You add the service's routes to the gateway's table and reference it for discovery, and you wire JWKS discovery so the new service resolves keys through that same gateway:
var gateway = builder.AddProject<Projects.MMCA_App_Gateway>("gateway")
.WithReference(engagementService)
.WithReference(conferenceService)
.WaitFor(engagementService)
.WaitFor(conferenceService);
engagementService.WithJwksDiscovery(identityService, gateway); // Profile A form (see below)
Centralizing the entry point keeps client config trivial (one URL), centralizes CORS and auth-forwarding, and lets the internal services run cleartext h2c without exposing that to clients (ADR-008).
Step 8: what the edge process itself owns
Step 7 gets a request to the right service. The edge process also has jobs of its own, and they are
jobs no service behind it can do, because only the gateway sees every request exactly once. Three of
them ship in the Gateway namespace of MMCA.Common.Aspire. A fourth, the forwarded-headers step
that makes the client IP the caller's rather than the ingress's, ships in MMCA.Common.Gateway
alongside AddMmcaGateway, the shared YARP forwarder profile both gateways chain onto
AddReverseProxy. The package placement is a constraint, not a preference: a YARP host has no
controllers, so it references those two framework packages and nothing else from the framework, and
it cannot reach the service-tier CorrelationIdMiddleware or AddCommonRateLimiting that live in
MMCA.Common.API.
// Gateway Program.cs
builder.Services.AddGatewayRateLimiting(builder.Configuration);
builder.Services.AddGatewayDownstreamHealthChecks("identity", "conference", "engagement", "notification");
builder.Services.AddReverseProxy()
.LoadFromConfig(builder.Configuration.GetSection("ReverseProxy"))
.AddMmcaGateway(builder.Configuration) // the shared forwarder profile and per-route policies
.AddServiceDiscoveryDestinationResolver();
...
app.UseCommonForwardedHeaders(); // so the client IP is the caller's, not the ingress's
app.UseGatewayCorrelation(); // before anything, including a 429, can short-circuit
app.UseCors();
app.UseGatewayRateLimiting(); // after CORS, before the proxy is mapped
app.MapReverseProxy();
- One correlation id for the whole hop, not one per service.
GatewayCorrelationMiddlewaredeclares theX-Correlation-IDconstant and, when the caller sent none, mints one fromActivity.Current?.TraceIdwithHttpContext.TraceIdentifieras the fallback. The part that makes it one id is that the value is written back onto the request headers before forwarding, so the downstream service's own middleware finds a header already present and adopts it instead of minting a second one. The response echo runs fromResponse.OnStarting, so it survives a proxied response whose headers the forwarder writes. The middleware's only constructor dependency is theRequestDelegate: no scoped service, noHttpContext.Items, which is what lets it drop into a host with no application container. - An edge limiter that counts anonymous callers.
AddGatewayRateLimitinginstalls a per-client-IP fixed window (120 requests per 60 seconds by default) chained with a process-wide concurrency cap (200 in flight), both rejecting with 429. The anonymous posture is the deliberate inverse of the service-tier limiter's: ADR-019 exempts anonymous traffic partly because public reads are served from the output cache, and that cache lives behind the proxy, so a flood is paid for in full at the edge before anything can be served cheaply. The two limiters answer different failures: the window answers one noisy source, the concurrency cap answers total in-flight work no matter how many sources produced it. Four kinds of request take the no-limiter partition on both limiters./health,/aliveand/.well-knownare always bypassed, because throttling them takes out the probes and JWKS discovery as collateral. A host adds its own prefixes throughBypassPathPrefixes, which is how Store exempts its Stripe webhook route and ADC its SignalR hub route. The other two are secret-proving headers, each off until its secret is configured:SyntheticTrafficSecret, so a load run measures the system rather than the limiter, andTrustedCallerSecret, for a server-rendered UI host whose every back-end call leaves from one address that the per-IP window would otherwise collapse into a single partition. ADC's gateway names the trusted-caller header in configuration and takes the secret itself from Key Vault. A request with no attributable client IP is not limited at all: that is fail-open, chosen over collapsing every unattributable caller into one shared bucket. The settings validate on both construction paths, the options pipeline (ValidateOnStart) and aValidator.ValidateObjectcall at registration, because the limiter closes over an eagerly bound copy and a caller can hand it settings without passing through options at all. - Readiness that reflects the downstreams, liveness that does not.
AddGatewayDownstreamHealthChecks("identity", "conference", ...)registers one check per named service, each GETting/alivethrough a service-discovery-resolved client under a 2 second budget, and tags themReady. The tag is the design: aReadycheck reaches/health/readyand never reaches/alive, so a downstream outage pulls the gateway out of the load balancer without restarting a gateway process that is perfectly healthy.
There is a fourth thing the edge deliberately does not do, and the refusal is the decision: no JWT pre-validation at the gateway. ADR-004 puts validation authority in the services, and a second validator does not add a check, it adds a second truth: two processes reading two JWKS caches can disagree across a key rotation, and the one that rejects is the one the caller sees. The saving would be one forward of a request the service was going to reject in microseconds anyway. The trigger to revisit is measured, not felt: when invalid-token traffic is a material share of forwarded volume, it becomes a cost argument, and the services keep validating either way (ADR-088).
The route table itself is configuration, not code. What step 7 calls "the gateway's table" is YARP
ReverseProxy configuration as the single source, 33 routes in ADC and 15 in Store, pinned in both
repositories by a RouteMapTests drift gate that drives each route through the real proxy pipeline
and compares the loaded IProxyConfig back against the pinned list in both directions. The
alternative shape, hand-written MapForwarder calls, states the same table three times inside one
repository (the registrations, the comment above them, and the list the test pins) with nothing
forcing the three to agree, which is how a live route ends up ungated. The per-destination HTTP
version policy from the next section lives in cluster configuration rather than in positional
arguments at a call site, so a host's transport profile is readable one line under the service it
describes (ADR-089).
A note you cannot skip: ADR-012's two Kestrel transport profiles
On a cleartext endpoint there is no TLS, so there is no ALPN to negotiate the protocol. Kestrel must be
told up front which protocols the cleartext port speaks, and the choice forces matching gateway-forward
and JWKS-discovery wiring. There are two coherent profiles, and picking the wrong one fails with
HTTP_1_1_REQUIRED on gRPC or a JWKS backchannel that cannot reach the auth endpoint.
- Profile A (serves inbound cleartext gRPC, including any bidirectional pair). Kestrel is
Http2-only on cleartext (h2c prior knowledge), so peer gRPC clients negotiate without TLS. The gateway must forward HTTP/2 (ForwardHttp2=true, andtransport: http2on the container ingress). JWKS uses the two-argumentWithJwksDiscovery(identity, gateway), because the HTTP/1.1 JwtBearer backchannel cannot reach an Http2-only endpoint directly, so it goes through the gateway, which terminates TLS and routes/.well-known/*on. Any service that serves gRPC over cleartext needs Profile A. ADC uses it. - Profile B (consumer-only, one-directional gRPC, gRPC rides the HTTPS/ALPN endpoint). Kestrel is
Http1AndHttp2; gRPC clients use the HTTPS endpoint where ALPN negotiates HTTP/2. The gateway forwards HTTP/1.1 (ForwardHttp2=false), and JWKS uses the single-argumentWithJwksDiscovery(identity).
A 2026 update is the cautionary tale here: Store originally chose Profile B because its gRPC edges looked
"consumer-only," but a one-directional topology still has services that serve inbound cleartext
gRPC, and Azure Container Apps cleartext ingress cannot deliver HTTP/2 to them under Profile B. The
result was HTTP_1_1_REQUIRED 500s in production. Store converged to Profile A. The lesson: if any
service in your fleet hosts an inbound gRPC server over cleartext, you are on Profile A, and you must
flip Kestrel, ForwardHttp2, the ingress transport, and the JWKS form together.
There is a sharper case the two-profile framing does not cover on its own: a host that needs both.
ADC's Notification service serves a SignalR hub (whose WebSocket transport needs the HTTP/1.1 Upgrade
handshake) AND an inbound live-channel gRPC ingress, and Store's Sales service runs the same
mixed-endpoint profile. Neither whole-host profile fits,
because the constraint was only ever per endpoint, not per host. The resolution is a mixed-endpoint
profile that applies both profiles inside one process: the default cleartext endpoint (port 8080) stays
Http1AndHttp2 (Profile B) so the WebSocket Upgrade still works, and a dedicated Http2-only grpc
endpoint (port 8081) serves the cleartext h2c gRPC (Profile A) for
LiveChannelPushService.PushToChannel, mapped as LiveChannelGrpcService. Peers resolve that ingress
through its named endpoint scheme http://_grpc.notification, not the default port, and in Azure
Container Apps it rides a dedicated internal TCP port mapping so one app can serve HTTP/1.1 WebSockets
and end-to-end HTTP/2 without hitting envoy's single-transport limit. The takeaway generalizes: when one
host must speak WebSockets and serve inbound cleartext gRPC, split the two protocols across two Kestrel
endpoints instead of forcing one whole-host profile (ADR-012's mixed-endpoint amendment).
The whole point: it stays reversible
Walk the steps back and you see why this is not a one-way door. Because transport lives at the edge and
the core talks to abstractions, a service can be re-collapsed into a combined host by changing
configuration, not code. Enable the peer module in the same host, drop the services.Replace(...) that
swapped in the gRPC adapter, and the in-process implementation resolves again. Flip MessageBusSettings
back to InProcessMessageBus and events dispatch in-process. The ModuleLoader boots the same module
code either way. That reversibility is insurance: a small team can adopt microservices without betting
that the split was correct on the first try.
Trade-offs, honestly
- Operational complexity multiplies. You now run multiple deployables plus a gateway, service discovery, a broker, and per-service databases instead of one process. Aspire hides a lot of this locally; production needs the matching Bicep and ingress configuration.
- Distributed-systems semantics are now yours. Cross-service consistency is eventual (outbox plus integration events). There are no cross-service transactions and no cross-database foreign keys. Consumers must be idempotent because broker delivery is at-least-once.
- The transport profile leaks into hosting. ADR-012 is not optional reading. The Kestrel protocol choice, the gateway-forward mode, the ingress transport, and the JWKS form are a coupled set, and a half-configured set fails in ways that only show up in the cloud, not locally.
- Bidirectional pairs need deliberate startup handling. The no-
WaitFor-on-the-reverse-edge trick is necessary, and it relies on the resilience pipeline to absorb the transient startup errors. It works, but it is a thing you must know rather than discover. - The edge limiter's number is per replica, and changing it is a restart. Both edge limiters count in one process's memory, so the fleet ceiling is the configured value multiplied by the replica count, which rises exactly when the load that motivated the limit arrives. And because the limiter closes over an eagerly bound copy of its settings rather than resolving options per request, no reload reaches it: calling the limit "a configuration section" invites the assumption that it can be changed live, and it cannot (ADR-088).
None of these are reasons to keep everything in one process forever. They are the reasons to keep the split reversible, so you only pay the complexity for the modules that actually need it.
Apply this even without MMCA
The recipe ports to any stack:
- Make the core talk to abstractions and put a fitness test on it. A message-bus interface and plain service interfaces, with an architecture test that forbids transport types in your domain and application layers, so the property cannot rot.
- Keep the wire contract in a shared package that generates both server and client stubs, and ship an adapter that implements the same interface the in-process code already used. Extraction becomes a registration swap.
- Federate auth with JWKS, not a shared secret, so adding a service does not mean distributing a symmetric key.
- Front the services with a single gateway so clients see one URL and internal transport choices stay internal.
- Treat the Kestrel/gateway/ingress/JWKS transport choice as one coupled decision, and write down which profile you are on. The half-configured states fail in production, not in development.
The takeaway: the monolith-to-microservices split is a rewrite only if you let transport leak into your business logic. Hold the abstraction boundary from day one, enforce it with a test, and the split becomes a wiring change you can undo.
What we covered: the starting invariant (core talks to abstractions, enforced by
MicroserviceExtractionTests), defining a .Contracts project with auto-compiled protos and a
same-interface gRPC adapter, server-side AddGrpcServiceDefaults() with Result-over-the-wire,
client-side AddTypedGrpcClient<T>(serviceName) with JWT forwarding and Polly over h2c, flipping
IMessageBus to BrokerMessageBus via settings, JWKS so the new service validates tokens, AppHost
wiring through MMCA.Common.Aspire.Hosting, routing through the YARP gateway, the edge
responsibilities the gateway process owns (forwarded headers, ensured correlation, an
anonymous-counting rate limiter, downstream-aware readiness) plus the JWT pre-validation it declines
and the route table it keeps in configuration, ADR-012's two Kestrel transport profiles, and why the
whole thing stays reversible.
Next in the series: resilience handlers and recovery objectives, the retry pipeline behind these calls plus the RTO/RPO targets and the drilled restore behind them.
MMCA.Common is Apache-2.0 licensed and open source. Star the repo, read ADRs 007, 008, 012, 088, and 089
behind this walkthrough, or dotnet add package MMCA.Common.Grpc and try it.
- ⭐ Repo: https://github.com/ivanball/MMCA.Common
- 📚 Full series index: https://ivanball.github.io/writing.html
- 📄 ADR-007 (gRPC extraction), ADR-008 (service-extraction topology), ADR-012 (gRPC-host transport), ADR-088 (gateway edge responsibilities), ADR-089 (gateway topology owned by configuration) in the repo.
Tags: .NET, C Sharp, Microservices, gRPC, Software Architecture