
Every guide to zero downtime database migration explains the same three phases and stops at the same place.
Expand the schema so old and new can coexist. Migrate the data and the traffic. Contract by removing the old path. Three boxes, one arrow between each, and the diagram makes it look linear.
The part that decides whether your migration goes well happens between boxes two and three, and it is this: there is a period during which rolling back is more dangerous than rolling forward. Two stores are both partially authoritative. The old one is missing everything written to the new one since the switch. Going back stops being a revert and becomes a second migration performed under pressure by people who have been awake too long.
Knowing when that window opens, keeping it short, and rehearsing the exit is most of the actual work.
For completeness, then we move on.
Expand. Add the new column, table or store alongside the old. Nothing reads it yet. All writes go to both, or writes go to the old and a capture process copies them across. The system is backward compatible, so any code version can run.
Migrate. Backfill historic data. Start shadow reads, where the new path is queried and its answer compared against the old one without being served. When the comparison is clean, shift read traffic gradually.
Contract. Stop writing to the old path. Remove it. Delete the column, drop the table, retire the adapter.
That is the whole pattern and it is correct. Everything that follows is what the diagram leaves out.

Draw a line at the moment the first write lands in the new store and not in the old one. That is where the window opens.
Before it, rollback is free. Turn off the new path, nothing is lost, go home.
After it, the old store is behind by everything written since. Rolling back means either accepting that data loss, which is usually unacceptable, or reverse migrating, which is the same work again in the opposite direction with an incident open.
Three design choices control that window.
Write to both, for as long as you can afford. Dual writing keeps the old store complete, so rollback remains cheap. It costs latency on every write and it introduces a partial failure case, which is the next point. But it is the single most effective way to keep the exit open.
Decide what happens when one write succeeds and the other fails. This is the question that separates a designed migration from an optimistic one. Fail the whole request, which makes your availability the product of both stores and can make things worse. Or accept the write, log the divergence, and reconcile asynchronously. The second is usually right, and it only works if the reconciliation actually exists, runs continuously, and alerts. A divergence log nobody reads is decoration.
Make the new path's writes idempotent and replayable. If a rollback becomes necessary, the ability to replay the divergence log into the old store turns a crisis into a script. Build that script during the expand phase, while nothing is on fire.

One practical rule that has saved more migrations than any pattern: do not start the window on a Friday, and do not start it before a holiday. It sounds unserious. It is the most reliable advice in this article, because the window's cost is measured in how quickly a decision can be made by people who understand the system.
The backfill is usually treated as a background chore. It sets everything.
Do the arithmetic before planning anything. A table of 40 million rows backfilled at 2,000 rows per second takes about 5.6 hours. The same table at 500 rows per second, which is what you get once you throttle to protect production, takes 22 hours. That is the difference between a single evening and a multi day operation with a handover.
Three things throttle you below the number you hoped for.
Replication lag. Writing hard to a primary pushes replicas behind, and read traffic served from a lagging replica returns stale data to users. Any backfill on a replicated system needs to watch lag and back off, which means throughput is variable rather than fixed.
Lock contention. Even online schema change tools take locks briefly, and on a hot table brief is enough to time out requests. Batch sizes that work in staging are frequently wrong in production because staging has no traffic.
The rows that fail. Some proportion will not migrate cleanly. Encoding, nulls where the new schema expects values, foreign keys pointing at rows deleted years ago. Expect between 0.1 and 2 per cent on a mature production table and design for it: failures go to a queue with a reason, the backfill continues, and somebody triages the queue. A backfill that halts on the first bad row will never finish.
Two properties make backfills survivable. Resumability, so a process killed at 60 per cent restarts from 60 rather than from zero. And rate control you can change at runtime, because the correct rate is discovered in production and you do not want a deploy to adjust it.
Shadow reads get recommended everywhere. The comparison strategy underneath them rarely does, and a weak comparison produces confidence rather than safety.

Read traffic should move in stages, controlled by a flag you can change without a deploy.
One per cent, then 10, then 50, then 100, with enough time at each step to see a full traffic cycle. A full cycle means a business day and a nightly batch, because the batch is where the unusual queries live and it is where migrations most often break.
Three details worth the effort.
Route by a stable key, not at random. Hash the account or user id so a given customer gets a consistent experience. Random routing means one customer sees the old data on one request and the new on the next, which produces support tickets nobody can reproduce.
Give yourself an instant, deploy free way back. The flag has to be changeable in seconds by whoever is on call, at three in the morning, without a pipeline. If reverting requires a deploy, you do not have a rollback, you have an intention.
Watch business metrics, not only technical ones. Latency and error rate will look fine while a subtly wrong result quietly changes an important number. Pick two or three business figures that should not move during the migration and watch those, because they are what will actually detect a correctness problem.
Here is the failure with the longest tail, and it does not look like a failure at the time.
Cutover succeeds. Reads come from the new store, everything is stable, and the team moves on to the next thing. The dual write is still running. The old column is still there. The adapter is still in the code path.
Two years later, that adapter is load bearing in a way nobody intended, the old column has drifted, and a new engineer spends two days working out which of the two is authoritative. The migration was never finished, it was abandoned at the point where it stopped being visible.
Contract properly and deliberately.
Scope drives most of the variance in these numbers, and two jobs get quoted alike when they should not be.
A schema change inside one datastore. Adding a column, splitting a table, changing a key. The transaction boundary still exists, so a write to old and new structures can be atomic. Consistency is largely free and the work concentrates in the backfill and the cutover.
A move between datastores. Postgres to a different engine, a monolith database split into services, a self managed instance to a managed one. Now there is no shared transaction. A write can succeed in one place and fail in the other, and no amount of care removes that case. It has to be designed for.
The second is roughly a third more work than the first for the same data volume, and the extra sits almost entirely in the dual write module, which is why that row carries the widest range in the table below.
There is a third case worth naming because it is often mistaken for the first. A change that alters the meaning of existing data, rather than its shape. Splitting a single amount column into net and tax, converting stored local times to a fixed zone, normalising a free text field into a reference. Historic rows cannot always be converted correctly, because the information needed was never recorded. That is not a migration problem with a technical fix. It is a decision about what to do with the ambiguous rows, it has to be made by someone from the business, and it should be made in week one rather than discovered at 70 per cent through the backfill.
Ask early which of the three you have. The answer changes the estimate more than the row count does.
The rate is $40 to $100 per hour by role. Tooling and cleanup sit near the floor, dual write reconciliation and cutover design near the ceiling, and mixed teams blend to around $65.
|
Module |
Hours |
What it covers |
|
Schema expand and compatible writes |
50 to 110 |
New structure alongside old, every code version able to run |
|
Dual write and reconciliation |
70 to 150 |
Both stores written, divergence logged, continuous reconciliation, replay script |
|
Backfill, throttled and resumable |
60 to 130 |
Rate control at runtime, resume from position, failure queue with reasons |
|
Shadow reads and comparison |
50 to 110 |
Whole response compared, sampled by shape, mismatches categorised |
|
Cutover and rollback rehearsal |
60 to 130 |
Staged flag, stable key routing, business metric watch, a rehearsed way back |
|
Contract and cleanup |
40 to 90 |
Dual write off, wait, read path, adapter, data, in that order |
Worked example. Expand 80 plus dual write 115 plus backfill 100 plus shadow 80 plus cutover 95 plus contract 65 equals 535 hours. That is $21,400 at $40, $53,500 at $100, and about $34,775 at a $65 blend.
The full span across the six modules runs 330 hours at every minimum to 720 at every maximum.
Deliberately excluded: the infrastructure for the new store, and any licence. Both are yours to buy and neither belongs inside an engineering estimate.
Weeks 1 to 2. Expand. Nothing reads the new structure and every deployed version still works.
Weeks 2 to 4. Dual write with reconciliation, plus the replay script that would undo it. Write the script now. It is much harder to write during an incident and it is the thing that keeps the exit cheap.
Weeks 4 to 6. Backfill in a staging environment sized like production, to discover the real throughput rather than the hoped for one. Then the production backfill, throttled, resumable, with the failure queue being triaged as it goes.
Weeks 6 to 8. Shadow reads. Do not shorten this. Every full traffic cycle you observe is a class of bug you will not meet at cutover.
Week 8. Rehearse the rollback on a copy. Actually do it, with the on call engineer driving, at an inconvenient hour. A rehearsal that is comfortable was not a rehearsal.
Weeks 8 to 10. Staged cutover, 1 per cent to 100, with a full business day and a nightly batch at each step above 10.
Weeks 10 to 12. Contract, in the order above, booked as work.
Twelve weeks, roughly 535 hours, and the exciting part reduced to a flag change that somebody has already practised reversing.
The measure of a good migration is that nobody outside the team notices it happened, and the measure of a well run one is that the people inside the team were not particularly worried either. Both come from the same place, which is having rehearsed the way out before needing it.
Expand the schema so old and new structures coexist and every code version can run, migrate the data and shift read traffic gradually, then contract by removing the old path. The pattern is well established. The risk sits in the transitions rather than in the phases.
The moment the first write lands in the new store and not in the old one. From then, the old store is missing everything written since, so rolling back means either losing that data or performing a reverse migration during an incident.
Dual writing keeps the old store complete and therefore keeps rollback cheap, at the cost of latency and a partial failure case that has to be handled explicitly. Capture based approaches are less intrusive and lag slightly. Either works if the divergence is logged and reconciled continuously.
Do the division. Forty million rows at 2,000 rows per second is about 5.6 hours, and the same table at 500 rows per second once throttled is 22 hours. Replication lag, lock contention and the 0.1 to 2 per cent of rows that fail all push the real rate below the hoped for one.
By shape rather than at random. Random sampling is dominated by the common case, which is also the case most likely to be correct. Over sample the largest accounts, oldest records, unusual character sets and unusual relationship counts, and compare whole responses rather than a single field.
About 330 to 720 hours at $40 to $100 per hour by role. A typical migration lands near 535 hours, roughly $34,775 at a $65 blended rate. New infrastructure and licences sit outside that figure.
Because the contract phase becomes invisible once cutover succeeds. The dual write, the old column and the adapter stay for years, and the next engineer inherits two schemas with no record of which is authoritative. Book the contract phase as scheduled work with a name against it.

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.
Share your details and we will talk soon.
Be the first to access expert strategies, actionable tips, and the trends actually shaping the digital world. No fluff - just practical insights delivered straight to your inbox.
Dive into our blog and stay ahead of the curve with expert perspectives, future-ready trends, and tech tips written for decision-makers and doers alike.