A promo usage limit needs compare-and-set, not a read then a write
A promotion capped at one use per customer and a thousand overall is two counters, and the obvious way to enforce them is a read, a check against the limits, and a write of the incremented values. Now picture two payment confirmations for the same code arriving at the same moment. Both read the same number, both pass the check, both write. The thousand-and-first redemption exists, and the customer who was allowed one has two. The whole defect lives in the window between the read and the write, and no amount of validation inside that window closes it.
We have built payment services, hosted checkout pages and marketplace flows for most of a decade, and promotion logic is where money bugs are cheapest to introduce and most awkward to explain afterwards. A discount larger than the basket is at least visible in a single order. An over-redeemed cap stays invisible until finance counts. Our promotion engine, promocodes, treats the counter as the hard part rather than the arithmetic, and this is the mechanism it settled on.
The read is a fact about the past#
Here is the shape that loses. It uses our real types, and the package deliberately does not let you write it:
usage = store.usage_for(code.code, customer=customer) # UsageCounts(999, 0)
if code.check(customer=customer, usage=usage, total=total) is None:
store.increment(code.code, customer=customer) # no such method, on purposeusage_for answers a question about the past. By the time check has agreed that UsageCounts(999, 0) sits under a UsageLimits(total=1000, per_customer=1) cap, the number may already be 1000. The second webhook is not racing the first one's decision, it is racing its own read, and it wins, because an unconditional increment has no way to know what it was authorised against.
That is why the RedemptionStore protocol has no increment on it. The only write the store exposes already carries what it believed when you called it.
The write carries the counts it was validated against#
def record(
self,
code: str,
*,
customer: Optional[str] = None,
expected: UsageCounts,
idempotency_key: Optional[str] = None,
) -> Optional[Reservation]:expected is the counts the caller validated. The store increments only when the counts it finds are still exactly those, and returns None otherwise. In the reference implementation the whole check is one comparison:
current = self._counts(key, who)
if current != expected:
return None
self._totals[key] = current.total + 1Two details are worth pulling out. First, UsageCounts is compared as a value, so the global total and the per-customer count are one condition instead of two. A redemption by a completely different customer moves the global total and therefore invalidates your pending write, which is conservative and correct, because the global cap is the rule that was checked against that number. Second, None is not an error. It is the store saying nothing happened, which is the only honest thing it can say, since it has no idea whether the code is still usable at the new counts. Deciding that is the caller's job.
In the tests, that refusal shows up on its own, with no concurrency at all:
first = store.record("SAVE10", customer="alice", expected=UsageCounts(0, 0))
stale = store.record("SAVE10", customer="alice", expected=UsageCounts(0, 0))
assert first.usage == UsageCounts(1, 1)
assert stale is None
assert store.usage_for("SAVE10", customer="alice") == UsageCounts(1, 1)A second write against counts that have already been spent is exactly what the losing webhook performs. Here it costs you a two-line test instead of a production incident.
On conflict the rules run again, not just the write#
The tempting retry is to re-read the counts and write again. That is wrong in a way you only see at the boundary: the reason the counts moved may be the reason the code is now unusable. So redeem retries the whole decision rather than the increment.
for _ in range(attempts):
usage = store.usage_for(code.code, customer=customer)
code.validate(customer=customer, usage=usage, total=total, now=moment)
reservation = store.record(
code.code,
customer=customer,
expected=usage,
...
)Every pass re-reads usage_for and re-runs validate. If the redemption that beat us to it was the last one the cap allowed, the next pass does not reach the write at all: validate raises PromoCodeRejected with Rejection.EXHAUSTED, and the loser of the race is told it lost for the right reason. If there is still room, the loop writes against the fresh counts and succeeds. Contention inside the budget simply does not reach the caller:
def test_contention_within_the_budget_still_redeems():
promo = PromoCode("CAP5", TEN, limits=UsageLimits(total=5))
store = FlakyStore(InMemoryRedemptionStore(), refusals=2)
receipt = redeem(promo, store, eur("80.00"), now=START, attempts=3)
assert receipt.usage == UsageCounts(1, 0)
assert store.inner.usage_for("CAP5") == UsageCounts(1, 0)FlakyStore is four lines of delegation that refuse the first n writes and then behave, which is enough to test a compare-and-set loop deterministically without threads. We reach for that shape constantly, because concurrency bugs become testable the moment you can inject the refusal instead of trying to provoke it.
The budget is finite (DEFAULT_ATTEMPTS is 3), and when it runs out redeem raises RedemptionConflict, which is a RuntimeError and explicitly not a rejection. That distinction is the operational point of the whole design. A rejection is an answer for the buyer: the code expired, the cap is gone, this basket is too small. A conflict is an answer for us: the code is so contended that three honest attempts could not land a write. There is nothing to tell a buyer there, and nothing a buyer could do about it. A webhook handler should let that bubble up so the provider redelivers, and a sustained conflict rate on one code says something about that promotion rather than about that customer. Collapsing the two into one "could not redeem" response is how a hot code turns into a wave of support tickets about codes that were valid the whole time.
A repeated key returns the earlier reservation#
Retried webhooks are not an edge case, they are the normal operating mode of every payment provider worth integrating. So the third outcome is a replay. When record gets an idempotency_key it has already seen, it returns the counts stored for that key with replayed=True and increments nothing:
first = capture(promo, store, cart, now=START, idempotency_key="evt_1")
again = capture(promo, store, cart, now=START, idempotency_key="evt_1")
assert first.replayed is False
assert again.replayed is True
assert again.usage == first.usage
assert store.usage_for("SAVE10") == UsageCounts(1, 0)Notice what the replay returns: the earlier reservation, not a fresh success and not an error. The caller gets the same usage it got the first time, so an idempotent handler can render the same receipt twice, and the replayed flag is there for anyone who needs to know the difference (logs, reconciliation, a counter on redelivery rates).
The key only works if you enforce it where the increment is enforced. InMemoryRedemptionStore holds its _totals, _per_customer and _keys dictionaries under a single threading.Lock, which honours the protocol across threads in one process and nothing more; the class docstring says so rather than letting anyone find out in staging. For a real store the protocol's requirement is specific: back the key with a unique constraint written in the same transaction as the increment. A key checked in application code before the transaction is just another read-then-write race, with a longer window.
The counters move while the buyer is paying#
The race between two webhooks has a slower sibling, the race between the checkout page and the payment confirmation. A basket priced at 10% off can sit in a payment intent for minutes while somebody else takes the last redemption, or the window closes, or the basket shrinks below the code's minimum.
So pricing and spending are separate calls. quote validates and prices without touching the counters. capture re-runs every rule through redeem against the clock and the counters of the moment the money actually moves, optionally against a changed total, and refuses with CaptureRejected, a PromoCodeRejected that carries the stale Quote, so a checkout API can tell the buyer that the code expired while they were paying. On refusal it records nothing:
with pytest.raises(CaptureRejected) as refused:
capture(promo, store, cart, now=START)
assert refused.value.rejection is Rejection.EXHAUSTED
assert store.usage_for("LASTONE") == UsageCounts(1, 0)The counter is still 1. Losing the second check costs the promotion nothing, and that property is what makes it safe to run the check at all.
What it costs to run#
One extra read per attempt, and a write that can come back empty. In the uncontended case, which is almost every case, that is one read and one conditional write, no locks held across a network call, and no transaction kept open while a rule is evaluated. Under contention the cost is bounded by attempts, and the conflict test asserts that bound directly: store.writes == attempts. Nothing retries forever.
What it asks of the surrounding system is modest and not negotiable. The store's conditional write and its key uniqueness must sit in the same transaction as the increment. The idempotency key must be a fact the provider gives you and you persist, not something derived at handling time. And the three outcomes have to stay three: a receipt, a rejection the buyer can act on, and a conflict that belongs in your alerting. A cap enforced this way is a property of the write rather than a hope about timing, and that is probably the only version of it that survives two webhooks arriving together.