Shipmind Labs

Hidden is not denied: one row declaration, two checks

· 8 min read

A delivery backend looks like one product, and it answers to half a dozen audiences, most of them reading the same order table. The list endpoint filters it correctly (a courier sees the shipments assigned to them), and the detail endpoint, written later and under different pressure, restates the same rule as a hand-written ownership branch. That is two copies of one decision, and they drift the moment you edit either one.

We have shipped this shape more than once: marketplaces with role-specific apps for couriers, warehouse staff and store operations, compliance tooling where a moderator must see every record, and back-office panels sitting on top of both. The hard part comes from the number of audiences rather than from the check itself. Five actors reading the same table want five different sets of rows, so a rule written inside a view is right for one audience and wrong for the rest. We pulled the pattern out into role-scopes, and the rest of this post is the argument it encodes.

Two questions, two tables#

What may this actor do and which rows are theirs are two different questions, and a backend that answers only the first one hands every courier the whole fleet.

So we keep two read-only tables. ACTOR_PERMISSIONS answers whether an actor may act at all. ACTOR_SCOPES answers which rows are theirs, keyed by actor and by the resource half of a permission name, so resource_of("shipment.deliver") is "shipment". The key is the resource and not the permission, which means shipment.view and shipment.deliver narrow through the same slice, declared once:

python
ACTOR_SCOPES = {
    Actor.COURIER: {
        "order": Scope.owned("order.own_assignment", shipment__courier_id="courier_id"),
        "shipment": Scope.owned("shipment.own_assignment", courier_id="courier_id"),
    },
    Actor.WAREHOUSE: {
        "order": Scope.owned("order.own_store", store_id="store_id"),
        "inventory": Scope.owned("inventory.own_store", store_id="store_id"),
    },
    Actor.SUPPORT: {
        "order": Scope.EVERYTHING,
        "shipment": Scope.EVERYTHING,
    },
}

A slice has three shapes. Scope.EVERYTHING for the actors that answer for the whole business. Scope.owned() for a filter resolved from the acting principal, where each entry pairs a queryset lookup with the attribute of the principal that fills it. And nothing at all: an undeclared resource is Scope.NOTHING, so receiving reaches no orders because nobody wrote down that it should. Reach gets granted, not inherited.

The endpoint that passes for the wrong person#

Here is the version that ships on the first pass:

python
@api_view(["POST"])
def deliver(request, shipment_id):
    require(request.user.role, Permission.SHIPMENT_DELIVER)
    shipment = get_object_or_404(Shipment, pk=shipment_id)
    shipment.mark_delivered()

The check is real, and it passes for the wrong person. Every courier holds shipment.deliver (that is what makes them a courier), so courier 7 posting to shipment 8821, assigned to courier 12, marks somebody else's parcel delivered. The list endpoint was scoped correctly and never showed 8821 to courier 7, and that is what lets the bug survive review: hidden is not denied, and an id is an integer.

The tempting fix is a branch in the view, if shipment.courier_id != request.user.courier_id. It is right for couriers and wrong for the other five audiences. Support reaches every shipment and would be locked out. Warehouse owns by store_id. And a courier's reach over an order is not a column at all, it is order.shipment.courier_id. Written by hand, that comes to six branches per endpoint, drifting apart one endpoint at a time.

The fix is to ask the second question against the declaration the list already used:

python
@api_view(["POST"])
def deliver(request, shipment_id):
    shipment = get_object_or_404(Shipment, pk=shipment_id)
    require_object(request.user.role, Permission.SHIPMENT_DELIVER, shipment, request.user)
    shipment.mark_delivered()

One slice, two consumers#

Both sides read the same Scope, and both apply the permission check first.

scope_queryset() returns queryset.none() when the permission is denied or the slice is NOTHING, the queryset unchanged on EVERYTHING, and filter(**scope.filters(principal)) on OWN.

check_object() walks the identical slice, one lookup at a time:

python
if scope.kind is ScopeKind.EVERYTHING:
    return decision
if scope.kind is ScopeKind.NOTHING:
    return denied(f"no {resource} slice is declared for {decision.actor}")

expected = scope.filters(principal)
for lookup, _attribute in scope.lookups:
    value = _object_value(obj, lookup)
    if value is _MISSING:
        raise MissingObjectKey(
            f"scope {scope.label!r} needs {lookup!r} on the {resource}"
        )
    if value != expected[lookup]:
        return denied(f"this {resource} is outside {scope.label}")

What makes this work rather than merely rhyme is _object_value. It splits the lookup on "__" and follows the parts across relations the way the queryset would join them: the courier's order slice is declared as shipment__courier_id, the list turns that into a join, and the object check turns the same string into order.shipment.courier_id. Neither side owns the rule, so neither side can quietly disagree with the other. The row a list hides is the row a detail refuses.

A missing key is an error, not a wider result#

Denial is not the failure mode we care most about. What we care about is a check that silently answers a question nobody asked. A principal arrives without the attribute a slice filters on (a support agent record with no courier_id, a service account, a user mid-migration). Filter on courier_id=None and you either match the rows whose column is NULL or you empty the list entirely. One of those is a leak, the other is a ghost bug report. So the resolution refuses:

python
def filters(self, principal: Any) -> dict[str, Any]:
    """Resolve the declared lookups against ``principal``."""
    resolved: dict[str, Any] = {}
    for lookup, attribute in self.lookups:
        value = _principal_value(principal, attribute)
        if value is None:
            raise MissingScopeKey(
                f"scope {self.label!r} needs {attribute!r} on the principal"
            )
        resolved[lookup] = value
    return resolved

MissingScopeKey is a LookupError, raised from the same place on both paths, because check_object calls scope.filters(principal) too. A misconfigured principal then fails loudly in staging instead of handing you a plausible page.

The object side draws one more distinction, and we think it is worth the code it takes. An attribute that is not on the object at all raises MissingObjectKey: that is a declaration pointing at a field the model does not have, which is a bug in the matrix. A relation that exists but is empty, say an order with no shipment assigned yet, resolves to None, compares unequal, and gets denied with Rule.OBJECT_OWNED. Unassigned is a legitimate state of the data. A lookup that cannot resolve is not.

Denials are objects, not booleans. Each one carries the actor, the action, the rule that said no and a reason, with as_dict() for an error body, so the 403 the client sees and the line in the log are the same sentence.

What it costs to run#

Adding a seventh audience is a row in each table and no view changes, because no view spells a rule out for itself. The price is that the declarations are global data: they belong under test as a matrix rather than as a scatter of per-view assertions, and a reviewer has to read two tables to know what an actor reaches.

The DRF adapter answers both questions from the same declarations, has_permission on every request and has_object_permission on the detail routes, and a queryset mixin narrows lists through the identical slice. Unmapped actions get refused rather than guessed: list, retrieve and create fall back to <resource>.view and <resource>.create, while update, destroy and custom actions are absent on purpose, because the matrix names domain actions (shipment.reroute) rather than HTTP verbs. Worth knowing: DRF only calls has_object_permission on routes that go through get_object(), so a function view that loads a row by id itself still has to call require_object. Installing a permission class does not close the trap in this post.

On the Django side, one auth group per actor is filled from the matrix by an idempotent sync that replaces a group's permissions instead of adding to them, which is what makes a capability removed from the matrix leave the group as well.

The operational cost we accept knowingly is the relation walk. A list of orders narrows in SQL, while the detail check for a courier follows order.shipment.courier_id in Python, so the object wants to be loaded with select_related on the detail route or the check pays a query. That cost is small and it sits in plain sight where it happens, and we prefer it to six hand-written branches that agree with the list view only until someone edits one of them.

The rule we hold to is simple: if a check can be derived from a declaration, derive it. Two places that must stay in agreement will not.

Was this useful?

Building something similar?

or email hello@shipmindlabs.com