Shipmind Labs

A state machine that hands back a half-moved object is worse than one that refuses to move at all.

The usual shape of a transition API goes: validate, fire the hooks, mutate the entity, save. The problem is that the mutation and the hooks live in the same mutable object, so when an after hook raises, you are holding something that is neither the old state nor a committed new one. The exception propagates, the entity in memory has already changed, and whether that leaks to the database depends entirely on who remembered to wrap the call in a transaction.

We took a different route in our open-source project order-lifecycle. The transition resolves the row, runs the before hooks, builds the new machine with a dataclass replace, runs the after hooks, and only then returns it. The new machine is a separate value that lives inside the call and never escapes unless every phase completed. So a failure in an after hook leaves you in the source state with no history entry at all, not a rolled-back one, just none.

The error has to carry more than a message. When a hook raises, it surfaces as a single failure type carrying the phase it happened in, the original cause, and the list of hooks of that phase that already ran before the raise. That last field is probably the one people forget. If three after hooks fire and the second one dies, you need to know the first one already sent its side effect, because that is what you compensate for.

Two supporting decisions fall out of this.

History is a frozen value. Recording an entry or a log line returns a new object instead of appending in place, so the half-written history of a failed transition cannot exist by construction. Entries default their timestamp to UTC now, but the transition call accepts an explicit clock, because replaying a backlog of events with wall-clock time stamps gives you a history that is ordered wrongly and silently.

Refusals are ordered and complete. Resolution rejects in a fixed sequence: an illegal transition first, carrying the triggers it would have allowed; then a role that is not permitted, carrying the required roles; then conditions that were not met, carrying every failed condition, because condition evaluation deliberately does not short-circuit. One round trip tells you everything wrong with the attempt instead of one problem at a time.

The pattern underneath all three is that the failure path needs to be designed with the same care as the success path, and it needs to be as informative. Most production incidents we have cleaned up in payment and compliance flows were not caused by a transition that failed. They were caused by a transition that failed halfway and told nobody which half.

If you run a lifecycle with side effects attached to transitions (notifications, ledger writes, document generation), the hook that succeeded before the one that died is the part you are left holding, and the error object is where we would rather see it named.

Was this useful?

Building something similar?

or email hello@shipmindlabs.com