↑↓ to navigate Enter to open "…" all these words ANDOR to combine

Auth & the edge · No. 22

Ephemeral by design: sub-second live channels over one SignalR hub

Conference-day polls and Q&A need small events fanned out to whoever is watching the page right now, and persisting them is wrong. Here is how MMCA.Common adds interest-group channels to its single SignalR notification hub and splits durable from ephemeral at the publisher boundary, without a second WebSocket.


You have a durable notification pipeline. It is careful and correct: it resolves recipients, writes one inbox row per user, and fires a best-effort real-time push, with the inbox as the source of truth. It is the pattern from Article 21 in this series, and it is the right shape when a user must be able to find the message minutes later.

Now the conference floor opens a live poll. Two hundred phones are looking at the same question, votes are landing several per second, and every one of them needs to light up the running tally on every other screen in under a second. Reach for the durable pipeline here and you inherit all the wrong costs. The outbox deliberately delays eligibility (Outbox:ProcessingDelaySeconds, default 5s, ADR-003) so an in-process handler can run first, which is exactly the latency you cannot pay. You would write a per-user inbox row for a tally that is worthless a second later and that nobody will ever open. You would persist data whose entire value is that it is on screen right now.

The other instinct is to stand up a second SignalR hub for live traffic. That doubles the WebSocket connection count per client and forces you to solve auth, reconnect, and the scale-out backplane twice, then split the client-side connection management in two. You now maintain two of everything to carry one more kind of message.

Two message shapes, not two transports

The mistake in both approaches is treating "durable notification" and "live update" as if the difference lives in the transport. It does not. Both are small messages pushed over a WebSocket. The difference is the delivery contract:

  • A durable notification is something a user must be able to find later. It is inbox-backed, it survives an offline user, and losing it is a bug.
  • A live channel event is something that only matters while it is on screen. It is broadcast to an interest group rather than a user list, it is never persisted, and it carries no delivery guarantee. Losing it is fine, because the next fetch shows the truth anyway.

ADR-039 makes that the whole design: one realtime transport, and the durable-versus-ephemeral decision is taken at the publisher boundary, not the transport. IPushNotificationSender for the first kind, ILiveChannelPublisher for the second. Same hub, same connection, two contracts.

The MMCA answer: one hub, two publisher paths

The single NotificationHub gains channel semantics. A channel is just a SignalR group named by a key like event:1 or session:123, and the hub gets its first client-invokable methods: JoinChannelAsync and LeaveChannelAsync (surfaced as the JoinChannel and LeaveChannel hub-method-name constants) map the calling connection into or out of that group. Alongside the existing ReceiveNotification method name it now also declares ReceiveChannelEvent, the client-side method channel events arrive on. Four method-name constants, one hub.

Publishing rides a new Application-layer port, ILiveChannelPublisher, that sits right beside IPushNotificationSender:

// The publisher interface. Application code depends only on the port, never on SignalR:
//   ILiveChannelPublisher.PublishAsync(channelKey, eventName, payloadJson, ct)   // ephemeral, no persistence
//   IPushNotificationSender.SendToUsersAsync(...)                                // durable, inbox-backed

// 1. The command handler commits the durable truth and returns the caller's fresh tallies.
//    It holds neither the publisher nor the queue: the aggregate raised LivePollVoteChanged,
//    and the broadcast hangs off that event instead.
await unitOfWork.SaveChangesAsync(ct);
var results = await resultsBuilder.BuildAsync(poll, command.UserId, ct);  // includes the caller's own vote
return Result.Success(results);

// 2. The domain-event handler runs after that commit, off the request path. It rebuilds the
//    tallies with no caller identity, so MyVoteOptionId stays null and no per-user field can
//    ride the broadcast.
var results = await resultsBuilder.BuildAsync(poll, userId: null, ct);
var channelKey = poll.SessionId is { } sessionId
    ? LivePollChannel.ForSession(sessionId)                               // "session:123"
    : LivePollChannel.ForEvent(poll.EventId);                             // "event:1"

// Enqueue returns void, deliberately: the channel is DropOldest, so it evicts to make room
// rather than refusing the write, and there is nothing here for a caller to branch on. Drops
// surface through the channel's itemDropped callback: DroppedCount plus a warning per drop.
liveChannelPublishQueue.Enqueue(new LiveChannelPublishWorkItem(
    channelKey,
    LivePollChannel.PollResultsChanged,
    JsonSerializer.Serialize(results, JsonSerializerOptions.Web)));

Three things in that block are load-bearing. First, the broadcast is enqueued after the commit, and it is the placement that guarantees it. Enqueue inline in the command handler and the enqueue runs while the transactional decorator still holds the transaction open, so a rollback can leave clients already told about a vote that never persisted. Hanging the enqueue off a handler for the LivePollVoteChanged domain event buys post-commit delivery with no sequencing code at all: in-process domain-event dispatch inside a transactional command is deferred until after the commit succeeds and dropped on rollback, so the guarantee follows from where the work is attached. The enqueue itself is a non-blocking write to a bounded in-process queue, and a single-reader hosted drain worker forwards each item to the publisher off the request path (BR-229), so a hub hiccup, a slow peer, or a down Notification service can neither stall nor fail the command.

Second, the payload is a pre-serialized JSON string. There is exactly one serialization point, at the enqueue edge, and the same opaque string rides every hop unchanged. The port's contract is PublishAsync(channelKey, eventName, payloadJson, ct), and no serializer type crosses it.

Third, a queue that evicts cannot report a drop through its return value. The channel is configured DropOldest, so the underlying TryWrite always succeeds: it makes room rather than refusing the write. A bool return would therefore be a promise the queue cannot break, and an "if the enqueue failed, log it" branch written against it is dead code that makes every drop under backpressure look impossible. So the port method is void Enqueue(LiveChannelPublishWorkItem workItem), and its own contract says the queue never rejects an item, leaving nothing for a caller to branch on. Drops go through the channel's itemDropped callback instead: each one increments a DroppedCount and logs a warning with the running total, so a drain falling behind is visible rather than inferred from missing client updates.

The implementation is chosen at the composition root, the same null-object discipline the durable sender uses. Infrastructure registers a no-op NullLiveChannelPublisher as the default, so the port resolves in every host whether or not real-time transport is configured. AddPushNotifications(configuration) then replaces it with SignalRLiveChannelPublisher, which is a nine-line class: it group-sends ReceiveChannelEvent through IHubContext<NotificationHub>. It never holds a hub connection itself, so a host that never calls AddPushNotifications keeps the no-op and the publishing code still resolves and runs with nothing on the wire.

Joining a channel, and why the key is validated

Channel membership is a property of the one existing connection, which is the point of not adding a second hub. On the browser side, NotificationHubService carries both traffic kinds on that connection. JoinChannelAsync and LeaveChannelAsync track membership, and OnChannelEvent registers multicast subscriptions, so an invisible layout listener and the page in front of you can observe the same channel concurrently without fighting over one callback. Every tracked channel is re-joined automatically on Reconnected, because SignalR group membership does not survive an automatic reconnect: the server starts the new connection with no groups, so the client has to rejoin or it goes silent after the first blip.

The join is not a free-for-all. JoinChannelAsync on the hub validates the requested key against PushNotificationSettings.ChannelKeyPattern (default NotificationScopeKey.Pattern, which is ^(event|session):[0-9]+$) and throws a HubException on a miss, so a client cannot subscribe itself to admin:secrets by inventing a group name. The pattern is a configurable regex (compiled once and cached, with a one-second match timeout), so the framework closes the obvious abuse vector without hardcoding any app's channel vocabulary.

Shape is not entitlement, though, and the hub says so in its own comment: a well-formed key is any authenticated caller's for the asking unless the host supplies one more thing. That thing is IChannelJoinAuthorizer, an optional, defaulted constructor dependency the hub consults after the pattern check and before the group membership exists, so a caller who is not entitled to a channel never receives a single payload published to it. The default of "no authorizer" keeps an existing host working unchanged, and a host that publishes anything to a channel that is not public to every signed-in user is expected to register one. The connection itself is bounded on the same hub: PushNotificationSettings.MaxConnectionsPerUser (default 20) caps live connections per user identifier per replica, and a refused connection hands its slot straight back, so aborted connections draining out cannot lock an account out of its own tab.

Crossing a service boundary: the gRPC ingress

Here is where the publisher abstraction earns its keep. In ADC the module that raises live events (Engagement, running the poll and Q&A handlers) is not the host that maps the hub. The Notification service maps it, because it is the only process whose IHubContext can reach the connected clients. So Engagement cannot call SignalRLiveChannelPublisher directly, there is no hub in its process.

ADR-039 anticipated exactly this: "a host that does not map the hub can replace the registration with its own transport." ADC does that with a gRPC adapter. Engagement calls AddNotificationLiveChannelClient(), which registers a typed gRPC client and Replaces (not TryAdds) the ILiveChannelPublisher registration with LiveChannelPublisherGrpcAdapter. The publishing handlers do not resolve that port at all: the vote and upvote domain-event handlers enqueue the pre-serialized payload onto the in-process publish queue and return. The single-reader drain worker resolves ILiveChannelPublisher per item, which in this process is the gRPC adapter, and it forwards the payload over PushToChannel under a tight two-second deadline, swallowing every failure (transport, resolution, broken circuit). Best-effort all the way down, and because the forwarding runs off the request path, neither a slow nor a down Notification service can add a millisecond to an Engagement vote.

The queue is applied everywhere, not just on the two hot paths. The four lower-frequency command handlers (closing a poll, opening one, submitting a question, moderating one) publish one broadcast per operator action, which is little enough to argue for awaiting PublishAsync inline. They enqueue anyway, and the argument for that is not latency: an inline await also makes every one of those commands hostage to a peer it has no reason to depend on. All four inject the queue and enqueue after the save, so not one command handler in the module awaits a gRPC publish, and none of them wraps a publish in a hand-rolled try/catch. The two question handlers do carry a guard, because their pending branch reads a fresh count from the database before it enqueues and that read must never fail a question that has already committed, but the guard is the framework's BestEffort.ExecuteAsync helper rather than a local catch: one Warning plus one increment of a besteffort.dispatch.failed meter tagged by operation name. What is guarded is a database read rather than a publish, which is a much smaller thing to have to swallow, and a broadcast path that has quietly stopped working is a number an operator can alert on instead of a log line nobody reads.

On the Notification side, LiveChannelGrpcService receives the RPC and simply delegates to its own ILiveChannelPublisher, which in that host is the real SignalRLiveChannelPublisher. The payload's opacity is what makes this trivial: no service on the path deserializes it, so no shared payload type has to cross the wire.

The one piece of real infrastructure this needs is the transport profile from ADR-012. Notification's default Kestrel endpoint stays Http1AndHttp2 so the SignalR WebSocket upgrade handshake still works, and its cleartext gRPC ingress lives on a dedicated Http2-only endpoint named grpc alongside it. That mixed-endpoint profile is what lets one service serve both a WebSocket to browsers and an h2c gRPC ingress to peers.

The queue is where you say what the work is worth

The publish queue in this article is one instance of a general contract (ADR-052): work that outlives a request goes on a bounded Channel<T> singleton, drained by a SingleReader hosted worker. What makes the contract interesting is that the same shape encodes two opposite correctness decisions, and the field that carries the decision is FullMode.

The live-channel queue is the permissive end:

private const int Capacity = 1024;

_channel = Channel.CreateBounded<LiveChannelPublishWorkItem>(
    new BoundedChannelOptions(Capacity)
    {
        FullMode = BoundedChannelFullMode.DropOldest,
        SingleReader = true,
        SingleWriter = false,
    },
    itemDropped: OnItemDropped);

DropOldest is the right answer here precisely because the payload is ephemeral. If the drain falls behind, the oldest pending broadcast is the least valuable thing in the queue, since a newer poll tally supersedes it anyway. A consequence worth knowing: under DropOldest, TryWrite always succeeds, because the channel evicts to make room rather than refusing. A caller checking the return value learns nothing, so a drop is observable only through the itemDropped callback and the discarded-broadcast counter it feeds.

The other end of the same field is FullMode = BoundedChannelFullMode.Wait, paired with the non-blocking TryWrite rather than an awaited WriteAsync. That combination makes a full queue refuse the request outright instead of either discarding it or blocking the request thread, which is the right answer when silently dropping work someone asked for is a bug rather than backpressure. A small capacity goes with it, because the bound then exists to refuse a runaway caller rather than to absorb a burst.

There is a prior question, though, and it decides whether a channel is the right home at all. An in-process queue is exactly as durable as the replica holding it: a deploy or a crash between the enqueue and the drain takes the pending items with it. That is the correct trade for a poll tally, whose value expires in a second anyway, and it is why the live-publish queue is the only bounded Channel<T> in MMCA.Common, MMCA.Store and MMCA.ADC. Work that a restart must not lose goes in a row instead. ADC's AI scoring pass over an event's sessions is that kind of work, so the organizer's request writes one durable internal command row and returns, and the framework's processor claims it, restores the requesting organizer's principal and runs the pass through the ordinary CQRS pipeline under a per-event distributed-lock claim taken with a zero wait, so a duplicate trigger skips rather than queues behind the run in flight. That mechanism, and how it sits beside channels and cron, is Article 51 in this series, "Four ways to do work later: channels, cron and durable internal commands".

The reusable idea is that "put it on a queue" is not a design decision, it is the start of one. The decisions are what happens when the queue is full, and whether losing the item on restart is acceptable at all, and the honest answers differ for a presence ping and for a job someone is waiting on.

The sibling contract: work a clock owns

The queue above and the durable row beside it answer the same question: a request started work and must not wait for it. Neither can answer the other one. Nothing starts a nightly retention purge except the calendar, and an in-process bounded queue is exactly the wrong home for it. It is empty at boot, it belongs to one replica, and it is gone on restart.

The first instinct is a periodic hosted service with a 24-hour interval, and it fails twice. An interval is not a time of day, so "every 24 hours" lands at whatever hour the last deploy happened to be and drifts with every restart after. And an interval-driven hosted service runs on every replica, so scaling a service to three instances silently triples the purge. The second instinct is Hangfire or Quartz.NET: a schema, a storage abstraction, a dashboard to authorize and host, and an upgrade obligation that every extracted service host inherits. ADR-074 declines both, on the grounds that the framework already ran a durable, multi-replica-safe polling loop in production. The outbox claims a batch of rows with one ExecuteUpdateAsync that stamps a lease and a token, and only one racing replica matches the predicate. The missing piece was a cron expression, not a product.

So the scheduler is that same claim-lease idiom pointed at a calendar:

// A job is a name, a schedule and a body. Nothing else.
public interface IScheduledJob
{
    string Name { get; }              // primary key of the persisted row: renaming it strands the old schedule
    string CronExpression { get; }    // five-field, UTC, parsed by Cronos; Scheduler:Jobs:{Name}:Cron overrides it
    Task ExecuteAsync(CancellationToken cancellationToken);
}

// The claim is what makes scale-out safe: two replicas racing on the same row both issue this
// update, and exactly one of them matches the still-unleased predicate.
var claimed = await context.Set<ScheduledJobEntry>()
    .Where(e => e.JobName == jobName
        && e.NextRunOn <= now
        && (e.LockedUntil == null || e.LockedUntil < now))
    .ExecuteUpdateAsync(
        s => s.SetProperty(e => e.LockedUntil, leaseUntil)
              .SetProperty(e => e.LockToken, lockToken),
        cancellationToken);

Three members, and the third one is ordinary application code. Jobs are resolved scoped, in a fresh DI scope per execution, the same way a request handler is, so a job body can take a unit of work, a repository or a command handler and nothing about it knows it is on a schedule. It also means the long-lived runner never captures a scoped dependency, which is the failure mode a hand-rolled timer usually ships with.

The state is one row per job (JobName as the primary key, plus the cron expression, NextRunOn, LastRunOn, the last outcome and duration, and the lease pair), and it lives in the Default data source only. That is a deliberate split from the outbox, which exists once per physical database because an outbox row has to be written in the same transaction as the aggregate that produced it. A schedule has no such tie: a job belongs to the host that registered it, so a four-database host gets one schedule rather than four copies of it contending for one occurrence. The token that won the claim also guards the outcome stamp, so a replica whose lease expired mid-run writes nothing and drops its stale result instead of overwriting the current holder's.

The runner is a plain BackgroundService rather than a fixed-period one, for the same reason the outbox is: after each cycle it sleeps until the earliest NextRunOn across the store, read through TimeProvider, capped at the polling interval (30 seconds by default) and floored at one second so an overdue row another replica already holds cannot spin the loop. A host with one nightly job is not waking 2,880 times a day to find nothing due, and because every timestamp comes from TimeProvider, a test can drive months of schedule in microseconds. Cronos (MIT, zero dependencies) is the one piece bought rather than built: it turns a string into the next occurrence and has no opinion about storage or hosting.

Two consequences worth stating plainly. A missed schedule runs once and then advances: the next occurrence is computed from the instant the run finished, not from the occurrence that was missed, so a host that was down for six hours fires one purge instead of six. That is right for a purge and wrong for a job that must produce an artifact per window, which has to make the window explicit in its own state rather than infer it from the schedule. And the lease is a time lease, not a fence: a run that overstays it can be claimed by another replica, so job bodies carry the same idempotency obligation the outbox already puts on event handlers.

The capability is opt-in, and the proof is a test rather than a promise: build the real model for a host that never called AddScheduledJobs and the job entity is simply absent, so a consumer that wants none of this gains no table in its next migration. In the adoption sweep all three Store services, three ADC services and the Helpdesk web host turned it on, each running its own runner against its own database. Engagement is the neat one for this article: the same composition root that swaps in the gRPC live-channel adapter registers the cron runner earlier in the same file. One host, three kinds of background work, and the thing that tells them apart is not the mechanism but who owns the work: a request, a clock, or a committed transaction.

Trade-offs, honestly

The ADR names the sharp edges, and they are all consequences of the ephemeral contract, not accidents:

  • Ephemeral means lossy. A client that connects after an event was published never sees it. This is the defining constraint, and it shapes every consumer: features must treat a channel event as a cache-invalidation hint over fetchable state, not as the state itself. The high-frequency tally events carry the fresh counts in their payload so a page can patch in place, but the safety net is always that the next fetch shows the truth.
  • Type safety is by convention. Handlers receive raw JSON and deserialize themselves. There is no compile-time contract on the payload; the discipline is a shared payload record in the consuming app's Shared project, referenced by both the publisher and the subscriber.
  • Durable notifications are still single-subscriber. Only channel events are multicast. The single settable NotificationCallback for durable notifications remains one subscriber. Unifying the two is deliberate future work, not something this decision blocks.
  • Multi-replica needs the backplane. If a hub-hosting service runs more than one replica, group sends only reach connections on other replicas through the Redis backplane that AddPushNotifications wires when a redis connection string is present. Single-replica deployments need nothing extra.
  • Best-effort is not delivery. A publish can still be discarded at three points: the queue evicts its oldest item under sustained backpressure, the drain worker swallows a publish failure, and the gRPC adapter swallows a transport failure. None of them are silent now (the eviction is counted on DroppedCount and logged with a running total, the drain's swallow is the shared BestEffort helper so it is both logged and counted on a meter, and the adapter's is logged per occurrence), but that is observability, not delivery. It is the correct posture for a hint over durable state, and it still means "the event fired" is never a guarantee. If you ever need one, you are describing a durable notification, which is the other publisher path.

Apply this even without MMCA

The idea ports to any real-time stack:

  1. Split on the delivery contract, not the transport. "Must survive an offline user" and "only matters on screen" are different guarantees. Give each its own publisher abstraction and let one connection carry both, rather than standing up a second hub.
  2. Attach the broadcast where post-commit is structural, not where it reads well. The durable write is the truth; the broadcast is an optimization over it. If your framework already defers domain-event dispatch until after the transaction commits, hang the broadcast off the event rather than writing it inline in the command handler, where a later rollback can leave clients told about state that never landed. Then hand it to a background drain (or at least wrap and swallow), log the failure, move on.
  3. Serialize once, at the edge, and keep the payload opaque across hops. A pre-serialized string makes a cross-service relay trivial: no hop in the middle needs the payload type.
  4. Validate client-supplied group names against a pattern. The moment a client can name the group it joins, an unconstrained join is a subscription to anyone's data. A configurable regex closes it without hardcoding your channel vocabulary.
  5. If a queue cannot meaningfully reject, do not give it a return value to ignore. An evict-on-full queue accepts every write by design, so a caller's "if the enqueue failed, log it" branch is dead code and every drop is invisible. Deleting the branch is the small fix. The honest one is to take the return value off the method, so the next caller cannot write the branch again. Then count and log the drop where the eviction actually happens, and carry a running total so a drain falling behind shows up in logs rather than in confused users.
  6. Ask who owns the work before you pick the mechanism. Work a request started and can afford to lose belongs on an in-process queue. Work a request started and cannot afford to lose belongs in a row, because an in-memory queue is only as durable as the replica holding it. Work a clock owns and a restart must not skip belongs in a row too, with a claim lease so three replicas do not run it three times. If you already have a durable polling loop (an outbox, usually), the gap between it and a scheduler is a cron parser, not a scheduling product.

The takeaway: real-time features do not need a second WebSocket, they need a second contract. Decide durable-versus-ephemeral at the publisher boundary and one hub carries both.


What we covered: why the durable notification pipeline is the wrong shape for high-frequency live updates, how ADR-039 keeps one SignalR hub and splits durable from ephemeral at the ILiveChannelPublisher boundary (with a no-op default, a SignalR group-send implementation, and channel join/leave with validated keys), why the hot-path enqueue belongs on a post-commit domain-event handler rather than inline in the command, how an evict-on-full queue has to report its own drops, how a remote module enqueues into the hub's channels through a best-effort, off-request-path gRPC ingress on the ADR-012 mixed-endpoint profile, why an ephemeral event must be treated as a hint over fetchable state, and how the sibling contract for clock-owned work (ADR-074) gets durable, multi-replica-safe cron out of the outbox's claim lease plus a cron parser, with no scheduling product.

Next in the series: explicit, compile-time DTO mapping that retires reflection-based AutoMapper and keeps a bad request a 400, not a 500.

MMCA.Common is Apache-2.0 licensed and open source. Star the repo, read the 2-minute ADR-039 behind this pattern, or dotnet add package MMCA.Common.Infrastructure and try it.

Tags: .NET, C Sharp, SignalR, Real Time, Software Architecture