Syntax is the manageable part of a Python 2 to 3 migration. The week when both versions are running at once is the part that costs you.
Codemods and a test suite cover the code. Data in flight is what they leave alone.
During a rolling migration, old workers and new workers share infrastructure: the same queues, the same cache, the same session store. A task serialized by one interpreter gets picked up by the other, and a cached blob written yesterday gets read today by a different runtime. The failures land there, not in the files you changed but in the bytes crossing a process boundary.
The fix is boring, and it has to happen before the migration starts. Move every cross-process payload to an explicit, interpreter-neutral format, JSON over pickle for queue messages. Put a format version into the cache key namespace, so the two runtimes do not read each other's blobs. Once that holds, the cutover goes per worker instead of per release night, and a rollback is just scaling one deployment back down.
We plan for the overlap window rather than pretending it won't exist. Both paths run, both are monitored, and the switch moves service by service.
If you have done a runtime or framework migration on a live system, we would like to hear what actually broke: the rewritten code, or something that travelled between two processes. Probably the second one.