Proof & getting started · No. 36
Soft-delete vs the right to erasure: the GDPR conflict and the erasure pathway

Soft-delete is the right default for lifecycle and audit. It is also, by itself, illegal for personal data under GDPR. This is the story of scoring that conflict 1 out of 4 in public, then fixing it.
Here is a deletion model almost every serious .NET app converges on, and it is correct:
public void Delete() => IsDeleted = true; // never actually remove the row
Nothing is hard-deleted. A global query filter excludes IsDeleted = true rows from normal queries, so
deleted records vanish from the app, but the data stays in the table. This is exactly what you want for
audit (you can prove what existed and when), referential integrity (a deleted parent does not orphan
its children), and undelete (a fat-fingered delete is recoverable). MMCA.Common does the same, with one
refinement: AuditableBaseEntity.Delete() flips IsDeleted and returns a Result (idempotent, so a
double delete fails cleanly instead of silently), EF Core global query filters exclude the row, and the
documented invariant is "entities are never hard-deleted."
Now read that same model through a privacy lawyer's eyes. A user invokes their right to erasure (GDPR Article 17, the right to be forgotten; CCPA deletion). They want their personal data gone. You run your delete. The row, including their name, their email, every personal field, stays in the database indefinitely, merely hidden from the application. You have not erased anything. You have hidden it.
Soft-delete and right-to-erasure are in direct conflict. The mechanism that makes lifecycle correct is the same mechanism that makes privacy non-compliant. And if your published privacy policy promises "we delete your data within 30 days," soft-delete cannot honor it.
Why it matters, and why I am telling on myself
When I first graded MMCA.Common against an architecture rubric and committed the scorecard to the repo, the original single-axis snapshot scored Compliance, Privacy & Data Governance at the bottom: a 1 out of 4, the lowest category on that early scorecard. The evidence was blunt: soft-delete everywhere with no erasure or anonymization mechanism, processed outbox payloads retained forever, no PII classification or data-subject-request scaffolding. The exact GDPR/CCPA conflict the rubric names, sitting unaddressed in the framework's defaults.
I could have quietly designed around it before publishing. I did the opposite: I scored it low, wrote the gap down with its evidence, and shipped the scorecard with the red flag intact. The reason is the whole thesis of scoring yourself in public. The gaps are the roadmap. A category that scores 1 with honest evidence is worth more than a category that scores 3 because nobody looked hard. This article is the "then I fixed it" half of that story, and the fix is ADR-005. On today's two-axis rubric, with ADR-005 shipped, the same category now scores Maturity 3 / Implementation 8.
The MMCA answer: separate the two concerns, give each a mechanism
The wrong fix is to overload Delete() so it hard-deletes personal data. That would break audit,
undelete, and referential integrity all at once, trading one defect for three. ADR-005 instead
separates the two concerns, because they answer different questions:
- Soft-delete answers "is this record active?" It stays the default for lifecycle and state management: hide, retain, undelete. It is explicitly not a privacy mechanism.
- Erasure answers "has this person's data been removed?" It is an explicit, additive capability, built on a new extension point rather than bolted onto delete.
That erasure pathway is the IAnonymizable interface (in MMCA.Common.Domain.Interfaces). An aggregate
that holds personal data implements it. An application-layer erasure handler loads the aggregate, calls
Anonymize(), and saves. Anonymize() overwrites the personal fields in place, so the row, its
foreign keys, and its audit trail all survive. The record still exists for integrity and accountability;
the person is no longer in it. The operation is idempotent and returns a Result, matching the
framework's domain conventions, so a retried erasure request is a safe no-op success.
// The hook: aggregates that store personal data implement this.
public interface IAnonymizable
{
Result Anonymize(); // idempotent: a second call is a no-op success
}
// ADC's User aggregate (an IAnonymizable AuditableAggregateRootEntity) implements it.
// Anonymize() overwrites every [Pii] field in place, including a crafted
// deleted-{Id}@anonymized.invalid email that KEEPS the unique-email invariant
// holding across many erased accounts. Audit fields and foreign keys survive.
This is the entity-level half of the split. In a consumer app (ADC's User), Delete() flips
IsDeleted for lifecycle and Anonymize() scrubs the data for a right-to-be-forgotten request, and
they are deliberately separate operations. The delete handler calls both in one transaction to honor
the "erase within 30 days" promise, and Delete() raises a UserDeleted domain event rather than
touching credentials. Outstanding refresh sessions need no separate revocation: the refresh flow
re-fetches the account through the same soft-delete query filter, so the moment the erasure commits
there is no account left to mint a fresh access token for. (The access token the user is already
holding is cut off separately, at runtime, which is its own section below.) For the fields a service chooses to keep retrievable after
anonymization rather than overwrite, ADR-005 offers an AES-256-GCM EncryptedStringConverter, so
"keep this usable" and "protect it at rest" need not be in tension. ADC's User takes the simpler
route: Anonymize() overwrites every [Pii] field (Email, FirstName, LastName, and the
AvatarUrl added with the avatar feature) in place and keeps none, so it does not wire the converter
at all. The converter is a framework capability for the
service that needs it, not a default.
There is one more piece that turns this from a convention into a guarantee: an architecture fitness
test asserts that any entity declaring a [Pii] property also implements IAnonymizable. You cannot
tag a field as personal data and then forget to give it an erasure path, because the build fails if you
do. The classification ([Pii]) and the obligation (an anonymize path) are wired together and enforced
by an executable rule, not left to a code review to remember. That is the §34 "executable governance"
idea applied to privacy: the rule that protects data subjects is a test, not a paragraph in a wiki.
The [Pii] marker carries a second duty: log and telemetry redaction. PiiRedactor masks every
[Pii]-marked member with a [REDACTED] token before an entity that holds personal data is written to
a structured log, and a PiiErasureContractFitnessTests build gate runs a [Pii]-bearing sample
through both the redactor and Anonymize() end to end, so neither half of the [Pii] contract can be
advertised without being exercised. The framework now wires the redactor into one production path of
its own: the opt-in field-level audit trail, whose save-changes interceptor asks the redactor whether a
type carries personal data and, for the properties that do, writes the redaction token in place of both
the old and the new value of the recorded change. One honest caveat for logging: PiiRedactor is still
an opt-in utility you call at the log site, not an auto-wired Serilog/ILogger destructuring policy. It
protects the call sites that adopt it, not every log line by default, so wiring it in is still the
consumer's job.
The delete that really removes rows: restrict by default
Soft-delete keeps the row. Anonymization keeps the row and destroys the person inside it. Physical
erasure, a real DELETE that takes rows out of the table, is a third thing, and the database will do
it on your behalf unless you say otherwise. EF Core picks a delete behavior for any relationship
nobody configured: a required one cascades, so removing a parent removes its children in the
database, below the aggregate's invariants, below the soft-delete filter, and below the domain events
the rest of the system reacts to. That is a destructive default nobody wrote down, and it sits
underneath every privacy story above.
ADR-119 inverts it. RestrictDeleteByDefaultConvention
(Source/Core/MMCA.Common.Infrastructure/Persistence/Conventions/RestrictDeleteByDefaultConvention.cs:41)
is an IModelFinalizingConvention that makes DeleteBehavior.Restrict the default for every
relationship nobody configured, registered per engine by the shared ApplicationDbContext at model
finalization (Persistence/DbContexts/ApplicationDbContext.cs:394). A delete that would orphan rows
fails loudly instead of quietly taking the children with it, and a genuine cascade becomes a decision
somebody recorded with .OnDelete(DeleteBehavior.Cascade) in an entity configuration. Two things are
deliberately left alone: an ownership foreign key keeps the cascade EF requires, and any behavior a
configuration already chose is kept exactly as chosen.
// Registered per engine, applied when the model is finalized.
configurationBuilder.Conventions.Add(_ => new RestrictDeleteByDefaultConvention(DataSourceKey.Engine));
// Every foreign key leaves the finalized model stamped with where its behavior came from:
// "Explicit" = an entity configuration chose it
// "Convention" = this convention restricted it
// "Ownership" = an owned type's cascade, which EF owns
That stamp is the point as much as the restrict is. Each foreign key carries a
MMCA:DeleteBehaviorSource annotation (:47) holding Explicit (:50), Convention (:53) or
Ownership (:56), and DeleteBehaviorConventionTestsBase
(Source/Hosting/MMCA.Common.Testing.Architecture/Bases/Domain/DeleteBehaviorConventionTestsBase.cs:21)
reads exactly that annotation to assert every cascading relationship was opted into on purpose. The
convention is a no-op on Cosmos, which has no foreign key constraints to restrict.
For a privacy story the payoff is direct: with restrict as the default, physically removing personal
data is always an explicit decision with a name attached, never something that happens as a side
effect of a required foreign key. Erasure stays a deliberate operation (Anonymize()), lifecycle
stays a deliberate operation (Delete()), and row removal is the third deliberate operation rather
than the one the ORM chose for you.
The second hiding place: the outbox
There was a second source of retained personal data, and it is the kind of thing you only find when you
go looking honestly. The transactional outbox (ADR-003) writes an event row in the same transaction
as the state change, and those serialized event payloads can contain personal data. The outbox
processor only set ProcessedOn; nothing ever purged processed rows. ADR-003 itself admitted the table
"grows until cleaned up." So even after you anonymized a user's aggregate, their personal data could
still be sitting in old, processed outbox payloads forever.
ADR-005 closes that too: OutboxCleanupService purges processed outbox rows older than
Outbox:RetentionDays (default 7, set 0 to disable) across every relational data source. Bounded
retention without changing delivery semantics. It is worth flagging as a deliberate behavior change:
consumers upgrading the framework start purging processed outbox rows older than seven days unless they
opt out, which is the kind of thing that belongs in a CHANGELOG and an ADR rather than a surprise in
production.
The live session: cutting off a token mid-flight
Anonymizing the aggregate and committing the soft-delete shuts two doors: no future logins, and no
silent re-mint of a fresh access token, because the refresh flow re-fetches the account through the
same soft-delete query filter
(Source/Core/MMCA.Common.Application/Users/UseCases/DeleteUser/DeleteUserHandlerBase.cs:103-107).
There is a third door, though, and stateless auth
holds it open by design. Authentication is stateless JWT (ADR-004): every service validates an access
token by signature and expiry, with no per-request lookup against the account store. That is exactly
what makes it scale, and it is also why soft-deleting a user does not, on its own, stop a token that was
already issued. A bearer credential keeps passing validation until it expires on its own clock, which
can be minutes after the account was deactivated.
ADR-047 bounds that window. SoftDeletedUserMiddleware
(Source/Presentation/MMCA.Common.API/Middleware/SoftDeletedUserMiddleware.cs:31, business rule
BR-133 named in its class doc at :11) runs in the shared pipeline after authentication and before authorization.
That ordering is a declarative step list (ADR-079):
Startup/Pipeline/MiddlewarePipelineBuilder.cs registers UseAuthentication() at :117, the
SoftDeletedUserFilter step that adds the middleware at :137, and UseAuthorization() at :141.
So HttpContext.User is already populated and the check
gates every downstream endpoint. For an authenticated caller whose account has been soft-deleted, it
returns HTTP 401 mid-flight, before the endpoint runs. The account-status lookup goes through
ISoftDeletedUserValidator
(Source/Core/MMCA.Common.Application/Interfaces/Infrastructure/Auth/ISoftDeletedUserValidator.cs:7,
one IsUserSoftDeletedAsync method at :15), which each Identity module implements with a single
filter-bypassing existence query, and the boolean result is cached for roughly 30 seconds
(SoftDeletedUserCache.MarkerDuration, Source/Core/MMCA.Common.Application/Auth/SoftDeletedUserCache.cs:29,
consumed by the middleware at SoftDeletedUserMiddleware.cs:132). So a given user costs at most one status query per cache window per
cache scope, not one per request. The effect is a bounded revocation window, not instant
revocation: a deactivated user's still-valid access token stops working within about the cache duration
instead of at the token's own expiry.
Two honest edges. Anonymous requests pass straight through with no lookup
(SoftDeletedUserMiddleware.cs:65-73), so unauthenticated traffic pays nothing. And the validator is
resolved lazily (SoftDeletedUserMiddleware.cs:75) rather than injected as a parameter, so a host that
does not register one (a non-Identity extracted service, or MMCA.Helpdesk's single Tickets host) simply
no-ops on that check: Identity is the source of truth and already validated the token upstream. Lazy
resolution is what keeps one pipeline correct in both Identity-hosting and non-Identity hosts without a
per-host variant.
The other right: portability, and why it degrades instead of failing
Erasure has a sibling that lands in the same regulation and is usually built badly. The right to data portability means a user can ask for everything you hold about them, and once your modules own separate databases, that request is a fan-out across services rather than one query.
The shape MMCA.ADC uses starts where you would expect, with the fan-out itself living in the
framework. ExportUserDataHandlerBase assembles a shared UserDataExportDTO envelope; the
app fills in the two genuinely app-specific parts. ADC projects its own account fields into a
UserDataExportSubjectDTO, deliberately excluding credentials (password hash and salt, refresh token,
external-provider key), since a data export that ships a password hash has turned a privacy feature
into a breach. The cross-service personal data arrives beside it as sections: ADC registers one
IUserDataExportSection per peer, Engagement and Notification, and over a process boundary those
contracts are the same gRPC adapters everything else uses.
The interesting decision is what happens when a peer is down:
catch (Exception ex) when (ex is not OperationCanceledException)
{
// Best-effort: the contributor failed after whatever resilience pipeline it uses.
// Degrade the section instead of failing the whole export. The reason handed back is
// deliberately generic; the exception detail goes to the log, never to the subject.
UserUseCaseLog.ExportSectionUnavailable(logger, ex, sectionName, userId);
return new UserDataExportSectionDTO
{
SectionName = sectionName,
Available = false,
UnavailableReason = UserDataExportSectionDefaults.UnavailableReason,
};
}
A section degrades to Available = false and the export completes. One peer outage never fails the
whole export. That catch now lives once, in the base, wrapping every registered section, so a section a
consumer adds tomorrow inherits the behavior instead of re-implementing it (and the section
contributors deliberately catch nothing themselves).
That is worth sitting with, because the instinct runs the other way. A partial export feels wrong, and the tidy engineering answer is to fail the request so the user retries and gets everything. In practice that is the worse outcome: a user exercising a legal right against a four-service system would be blocked by any one service having a bad afternoon, and you have converted a routine availability blip into a compliance failure. Marking the section explicitly unavailable is more honest than both alternatives, because it neither pretends the data is absent nor blocks the parts that are ready.
Two details keep it honest rather than sloppy. OperationCanceledException is deliberately excluded
from the catch, so a genuine cancellation propagates instead of being silently recorded as a missing
section. And the degradation is logged at warning with the section name and user id, so "this export
was incomplete" is an operational event someone can find later, not a silent gap in a file the user
already downloaded.
The transferable rule: in a distributed system, "all or nothing" is a choice with a cost, and for a read-only aggregation the cost is usually higher than partial success. Make the partiality explicit in the payload and loud in the logs, and it stops being a lie.
Trade-offs, honestly
The framework provides the mechanisms (IAnonymizable, OutboxCleanupService); the consumer owns the
policy. That division is the most important caveat, and ADR-005 states it plainly.
- Erasure is opt-in per entity. An aggregate that holds personal data but does not implement
IAnonymizablewill not be erased. Consumers must audit their own personal-data inventory and implement the interface where it applies. The framework cannot find your PII for you. - The framework cannot make you compliant on its own. It ships the mechanism, not the obligation.
The consumer is the data controller: it still has to wire the erasure handler, the data-subject
request flow, and the access/export endpoint, because the personal-data model lives in the consumer
(ADC's
User, with[Pii]-tagged fields, and its ownUserDataExportSubjectDTOinside the shared export envelope, deliberately omitting secrets). This ADR provides the mechanisms, not the policy. - Anonymization is irreversible by design, and it is not undelete. Soft-delete is recoverable; erasure is not. Conflating them would be a bug. They are separate operations precisely because they must behave differently.
- The default 7-day outbox retention is a behavior change on upgrade, as noted above.
None of these are reasons to skip the mechanism. They are the boundary between what a framework can give you (a correct, audit-preserving mechanism) and what only you can decide (which data is personal, and what your privacy policy promised).
Apply this even without MMCA
The pattern is independent of the framework:
- Keep soft-delete for lifecycle. It is the right tool for "is this active?", and audit / undelete / referential integrity all depend on it. Do not throw it out.
- Do not overload delete to satisfy erasure. Hard-deleting personal data inside your soft-delete path breaks audit and integrity. Add a separate erasure operation.
- Anonymize in place. Overwrite personal fields with placeholders, keeping the row and its foreign keys. Craft replacement values that preserve invariants (a unique-email constraint needs a unique anonymized email per record). Make it idempotent so retried requests are safe.
- Hunt the second hiding place. Personal data leaks into event logs, outbox tables, message payloads, and caches. Give every store that can hold PII a bounded retention policy.
- Write down where the framework stops and your obligation starts. A library gives you the mechanism; classifying your PII and honoring your privacy policy is your job as the data controller.
The takeaway: soft-delete and erasure answer different questions, so give them different mechanisms. And if you are going to grade your own architecture, grade it honestly. The category I scored a 1 became the most useful entry on the scorecard, because it turned straight into ADR-005 and a real erasure pathway.
What we covered: why soft-delete (the right default for lifecycle and audit) directly conflicts
with the GDPR/CCPA right to erasure, why the original single-axis snapshot scored it at the bottom (1
out of 4, now Maturity 3 / Implementation 8 on the two-axis rubric), and how ADR-005 resolves it by
separating the concerns: an IAnonymizable anonymize-in-place hook
that preserves the audit trail, plus an OutboxCleanupService that bounds retention of PII-bearing
outbox payloads, with the consumer owning the policy. And how ADR-047's SoftDeletedUserMiddleware
closes the third door, cutting off an already-issued access token mid-flight within roughly a 30-second
cache window instead of letting it live to its own expiry.
Next in the series: the reusable Blazor UI framework that brings the same backend discipline to the front end, a server-paged list page in a few lines.
MMCA.Common is Apache-2.0 licensed and open source. Star the repo, read the 2-minute ADR-005 behind this decision, or read the §30 scorecard entry, the most honest one on the board.
- ⭐ Repo: https://github.com/ivanball/MMCA.Common
- 📄 ADR-005 (soft-delete vs right-to-erasure): ADR-005: Soft-Delete vs. Right-to-Erasure in the docs site.
Tags: .NET, C Sharp, Software Architecture, GDPR, Data Privacy