Shipmind Labs

An offline queue that keeps global order is a queue that stops working the first time one change fails.

That trade-off is the one most offline-first implementations get backwards. The instinct is to hold pending changes in one list and drain it in the order the user made them. It feels safe. It is also why a single rejected edit on one record can hold back thirty unrelated changes sitting behind it.

Order does matter, but it rarely matters globally. It matters per entity. If a courier renames a stop and then marks it delivered, those two have to land in that sequence. If some other stop was edited in between, it has no stake in either of them.

So in offlinequeue, our open-source library, order is scoped by entity. Each operation names its entity through a function we call entityOf, which defaults to the operation kind, and changes that share an entity form a lane. A lane is attempted oldest first and stops at the first change that has to be tried again, because anything after it may depend on it. Lanes do not wait for each other. A flush runs up to a concurrency limit of lanes at once, four by default.

That gives you two guarantees that usually fight each other: causality per entity, and the rest of the app carrying on syncing while one record is stuck. A change can also be parked, and its lane continues past it instead of treating one poison message as a full outage.

The second half is idempotency, and that is the part teams skip. Keys are generated on the device at enqueue time, not by the server. Enqueueing a key that is already in the queue returns the row that is already there instead of creating a second one. And the key does not change across retries, so the same change carries the same key on its first attempt and on its fifth.

That last detail is what makes the retry story honest. Without a stable device-side key, a timeout is unresolvable: the client cannot tell a request that never arrived from one that succeeded and lost its response, so it either risks a duplicate or risks dropping the change. With a stable key, a repeat on the wire is something the server can recognise and collapse. Retry safety comes from the key, not from the network code.

We built this after enough mobile and field-operations work to know how it fails in production. The bug reports are rarely "sync is slow", they are "the change disappeared" or "it saved twice", and both trace back to ordering and key decisions made early and then left alone.

If you run offline writes, it is worth being precise about what your unit of ordering actually is, and about what happens to the queue behind a change the server keeps refusing. We think those two answers describe your sync layer better than any benchmark does.

Was this useful?

Building something similar?

or email hello@shipmindlabs.com