Your audit trail is a second copy of every sensitive value that changed
An audit trail records which fields changed on a row, and that means it also records the values those fields held. If one of them is a password hash, a card number or a passport number, the trail has become a second copy of that value, living in a log sink, replicated into dashboards and backups, retained on a different schedule and readable by more people than the table it came from. The part of the trail worth keeping is the fact of the change and the name of the field. The value is the part that creates the liability.
The copy nobody provisioned#
On the systems we build most often (payment services, KYC/KYB moderation tooling, e-signature workflows, lending back offices) the sensitive columns are the ones under the tightest control. Access is narrow, admin screens are restricted, retention is deliberate, and somebody has signed off on all three.
Auditing usually arrives later, and for a good reason: a compliance team needs to know who moved an account between issuance tiers, or a support lead needs to see which operator edited a payout requisite. The records then go wherever the project already keeps its logs, which in our stacks is often an index read through Kibana. That index inherits none of the controls the column had. Access is granted per dashboard rather than per row, retention follows the log policy, and by then the records are JSON, which puts them outside the project's data model and outside its permission model with it.
That asymmetry is the whole argument for withholding values at the moment the record is built. By the time a record reaches a sink, the decision has already been made for us. So in model-audit (https://github.com/shipmindlabs/model-audit) withholding is a property of the library rather than a setting somebody remembers: the values of sensitive fields are replaced before a record leaves the process.
Redaction keeps the fact, withholds the value#
The decision is per field, and it is the same decision everywhere:
from model_audit import DEFAULT_REDACTOR, MISSING, REDACTED
DEFAULT_REDACTOR.value("name", "Ada") # 'Ada'
DEFAULT_REDACTOR.value("phone", "+31 6 1234 5678") # REDACTED
DEFAULT_REDACTOR.value("phone", MISSING) # MISSINGA changeset is walked field by field through that same rule:
from model_audit import Changeset, FieldChange
changeset = Changeset(
(FieldChange("name", "Ada", "Grace"), FieldChange("api_key", "a", "b"))
)
DEFAULT_REDACTOR.changeset(changeset).as_dict()
# {'name': ('Ada', 'Grace'), 'api_key': (REDACTED, REDACTED)}The record still names api_key and still says it changed. What the reader loses is the only thing they had no business reading.
REDACTED is a str subclass whose single instance reads [redacted]:
import json
lookalike = "[redacted]"
isinstance(REDACTED, str) # True
REDACTED == lookalike # True
lookalike is not REDACTED # True
json.dumps({"phone": REDACTED}) # '{"phone": "[redacted]"}'Two properties follow from that choice. Every receiver that already serializes a record keeps working, because no sink needs a special case for a sentinel type, and that matters since the sink is usually code we did not write. And because identity survives, value is REDACTED separates a value we withheld from a value somebody typed that merely looks like one. That distinction stops being academic the first time a free-text notes field contains the literal string and a downstream consumer has to decide whether to trust it.
An addition still has to read as an addition#
Redaction flattens values, so it can also flatten a distinction a reviewer genuinely needs: was this the first time a phone number was set, or was an existing one replaced? The diff core reports a field present on only one side with MISSING on the other, and FieldChange exposes that as is_addition and is_removal. Redaction leaves MISSING alone:
redacted = DEFAULT_REDACTOR.change(FieldChange("phone", MISSING, "+31 6 1234 5678"))
redacted.old is MISSING # True
redacted.new is REDACTED # True
redacted.is_addition # TrueIf both sides collapsed to REDACTED, every change to a sensitive field would read identically, and the trail would answer "something happened here" instead of "a number was added to an account that had none". The shape of a change is not sensitive. Only the value is.
Exclusion is a separate and deliberate choice#
Withholding a value and refusing to mention the field are different decisions, and the API keeps them apart:
from model_audit import register
register(Customer, redact=["nickname"]) # 'nickname changed', values withheld
register(Customer, exclude=["nickname"]) # no trace the field was touchedThe first records the change by field name with the values withheld. The second leaves nothing at all: somebody reading the trail cannot tell the field exists, let alone that a write touched it.
Most exclusions have nothing to do with secrecy, they are about noise. Fields that change on every save drown the ones that matter, so NOISY_FIELDS (modified, modified_at, updated, updated_at, last_seen) are excluded unless you ask for them with ignore_noisy=False, and a derived column is an obvious manual exclusion:
register(Invoice, exclude=["search_vector"])The same split exists in the framework-free core, where exclude is just a comparison filter:
from model_audit import diff
stored = {"title": "Draft", "views": 10, "published": False}
incoming = {"title": "Release notes", "views": 10, "published": True}
changeset = diff(stored, incoming, exclude=["views"])
changeset.fields # ('title', 'published')The question that decides between the two is not how sensitive the value is. It is whether anyone could ever need to know that the field changed. For a password or an API key the answer is yes, since a credential rotation nobody can account for is exactly the event an audit trail exists to surface, so those are redacted and never excluded. Exclusion is right when the change carries no information, and right in the rarer case where the fact of the change is itself sensitive. That second case deserves to be written down and argued for, not reached by habit, because it is the one choice that leaves a reviewer with no way to know what they are missing.
Matching names is a heuristic, so it has an escape hatch#
Sensitivity is decided by a substring test. Field names are normalized (lower-cased, trimmed, hyphens turned into underscores) and checked against SENSITIVE_FIELDS, so X-API-Key and " Authorization " both match. A substring test is crude on purpose, preferring to catch document_number through document rather than miss a variant nobody thought to list. The price is false positives, and a redactor that quietly hides a field a team needs is a defect of its own:
from model_audit import Redactor
redactor = Redactor(allow=["document_type"])
redactor.is_sensitive("document_type") # False
redactor.is_sensitive("document_number") # TrueExtending is immutable, so a per-model redactor cannot widen or narrow the default for the rest of the process:
extended = DEFAULT_REDACTOR.extend(["nickname"])
extended.is_sensitive("nickname") # True
DEFAULT_REDACTOR.is_sensitive("nickname") # Falseemail is not in the defaults, and that is a decision rather than an oversight. An email address is usually the identity the trail is read by, and a change to it is one of the clearest signals of an account takeover, so a trail that withholds it hides the thing it was opened for. A project that needs it withheld adds it in one call, which is the right way round: redacting it by default would break the common reading of the trail, while adding it is explicit and local.
The request side carries the same values in a different shape#
Model records are only half the exposure. The same card number arrives in a request body, and request logs are usually even easier to read than an audit index. Route auditing is declared rather than blanket:
audit_routes({"POST,PUT,PATCH,DELETE /api/invoices/**": True})
subscribe_requests(write)Captured bodies and query parameters go through the same redactor, which walks mappings and sequences recursively, the shape a JSON body has on its way to a sink. Headers pass through a redactor extended with header patterns, so an Authorization value is withheld while the header keeps its place in the record. Bodies the middleware will not copy are replaced by a marker instead of a guess: {"omitted": "body too large", "bytes": ...} for an oversized payload, {"omitted": "unsupported content type", ...} for anything that is neither JSON nor a form. A record that admits what it did not capture is more useful to you than one that is silently partial.
What it costs to run#
register() connects post_init and post_save for one model: a snapshot of the row as it was loaded, a diff against what was written, a record handed to the receivers. Models nobody registered emit nothing, and a path the route map does not match costs a lookup. Redaction itself is a substring scan over the names of the fields that changed, set against the cost of the database write that triggered it.
The real cost is editorial. Somebody has to keep the sensitive list honest as models grow, decide for each new field whether it is withheld or absent, and revisit the allow list when a false positive is released. That work does not disappear if it is skipped, it moves to whoever eventually reads the trail.
The library records and never stores: subscribe() hands each AuditRecord to the receivers, and where the trail lives and for how long stays the project's decision. That is precisely why withholding has to happen before a record is emitted. Once it is in the log pipeline we no longer control who reads it, and the only values that cannot leak from there are the ones that were never written into it.