Skip to content
  1. Home
  2. Articles
  3. Complete Is Not Correct: Data Validation and Cutover When Replacing a System
113
Enterprise

Complete Is Not Correct: Data Validation and Cutover When Replacing a System

Complete Is Not Correct: Data Validation and Cutover When Replacing a System

Migrating every table does not mean the data is right. How to validate in layers, run cutover from a runbook, and have rollback ready before go-live.

The night before an organization switches to a new system, the question in the room is rarely whether the migration is finished. The migration team has already reported that every table has moved and that row counts match between source and target. On paper, the job looks done. The question still open is a different one: if we go live tomorrow, what breaks?

"Complete" and "correct" are not the same thing, and the gap between them is where a new system tends to stumble in its first week. Data that arrived row for row can still be wrong, can still be unusable, and can still bring the business to a halt. This article is written from the side of the team that has to live with the system after go-live, not the one that closes the project and walks away. We will look at what has to happen between "the migration is done" and "we are confident enough to switch over", the shortest stretch of the project and the one that decides how the whole thing turns out.

Complete is not the same as correct

Matching row counts are the first check, not the last. Every row can be present while the values inside have drifted during conversion between the old and new systems. A date stored in one format may land in another. A figure with two decimal places may get rounded. Thai text stored in a legacy encoding may turn into unreadable characters in the new system. None of this changes the number of rows, so all of it passes a row count without a problem.

Proving that the data moved correctly takes several layers. The first is making the counts match: necessary, but it proves the least. The second is a checksum or hash across numeric fields, such as the sum of every transaction amount in an account. If source and target produce the same figure, the chance that values changed in transit drops sharply.

The layer teams most often skip is the spot-check: pulling real records and comparing them one by one against the old system, deliberately choosing edge cases such as negative values, records with empty fields, or the oldest records in the system. Conversion errors tend to surface at the edges before they show up in the middle. The last layer is referential integrity, confirming that the relationships between tables survived. Purchase orders should still point to customers that actually exist, rather than becoming orphaned records linked to data that disappeared during the move.

Missing data gets noticed. Wrong data does not.

That is why data that is complete but wrong is more dangerous than data that is visibly missing. A missing record gets chased down. A record with the wrong value keeps being used quietly, feeding reports and decisions, until someone stumbles over it months later.

Agree on the pass criteria before go-live day

Acceptance criteria are the conditions agreed in advance for what counts as a pass, and they have to be measurable. A sentence like "data is accurate and complete" sounds fine, but it can be read a hundred ways, and nobody can say with certainty when it has been met. Usable criteria are specific. The financial total for every account matches the old system down to the smallest currency unit. The number of active customers is identical. Every record that fails a check carries a written explanation, instead of being waved through as a residual balance nobody looks at.

Good criteria share one property: each can be answered yes or no, rather than with a feeling that things are probably fine. When every criterion comes back yes, going live becomes a decision backed by evidence.

Just as important as the criteria is who signs them off. That person should be the business owner of the data, someone who works with it every day, because they are the one who can tell whether the numbers on screen make sense. The technical team can confirm that the migration followed the process. It cannot confirm, on the business's behalf, that this month's sales figure in the new system is the right number. The Go/No-Go decision therefore needs a clear owner and written criteria, not the shared mood of whoever is still in the room at 2 a.m.

Cutover has steps you cannot undo, so it needs a runbook

Cutover is the short window in which you actually switch from the old system to the new one. It is often pictured as copying the final batch of data and flipping a switch. In practice it is a sequence of dozens of steps that must run in order, and some of them cannot be reversed once they have started.

The core of this window is the runbook, a document that lists every step in time order. Each line should state four things: what is done, who owns it, roughly how long it takes, and which step has to finish first. On the night itself, nobody should have to work out on the spot what comes next, because improvising at 3 a.m. is where mistakes come from.

The sequence usually includes a freeze window, a period when the old system temporarily stops accepting new data so that no transactions appear after the final batch has been moved. Choosing its length is a trade-off worth discussing openly. A longer freeze allows more thorough migration and checking, but the business is paused for longer. A shorter one raises the time pressure and the chance of error. That is why many organizations schedule cutover over a weekend or a long holiday, when transaction volume is usually low.

Before the real date, run at least one mock cutover: a full rehearsal of the sequence with real data. The point of the rehearsal is to time the whole process and see whether it fits inside the window you actually have. Proving that it works is a side benefit. Many projects only discover during the rehearsal that their sequence takes more than one night, and that is something to learn a week before go-live, not at 4 a.m. on the night itself.

The rollback plan exists before you start, not after something fails

Before cutover begins, the team should already know at what point it will fall back to the old system. Those conditions should be no-go triggers defined in advance as numbers or clear events. For example: if more records fail validation than the agreed threshold, or if the data load is still unfinished past a set time, roll back. Having the conditions in writing beforehand spares everyone an argument at the moment when they are exhausted and the incident is unfolding.

What makes rollback harder by the minute is time. The longer users have been entering new transactions on the new system, the more expensive it is to go back, because every transaction created after the switch has to be carried back into the old system, or that period's data is lost when you fall back. Google Cloud's guide Database migration: Concepts and principles (Part 2) (last updated April 29, 2025) describes a fallback as a migration in reverse, and notes that the applications that ran against the original databases must be kept available so they can be switched back on.

A good fallback plan therefore has a deadline of its own, a point of no return after which the only option left is to fix forward. The team should know in advance where that point sits, and agree on who has the authority to call a rollback before it is reached.

An unplanned outage after go-live is a different problem from a rollback prepared in advance, with different roles and different decisions. We cover who takes command and how to communicate when a system actually goes down in The DR Plan That Survives the Day.

Parallel run: when it pays off and when it does not

A parallel run means operating the old and new systems side by side for a period, feeding the same data into both and comparing the results. It gives clear evidence that the new system produces correct results on real work, because the old system is there as a constant reference. Every difference between the two is something to trace back to its cause, and those differences often expose bugs that earlier testing missed.

The downside has to be stated plainly. A parallel run puts a heavy load on the team for as long as it lasts, because every transaction has to be handled in two systems and the results compared daily. If your case is an internal system with a few dozen users, one that can tolerate a short pause and holds no financial data that must never be wrong, a full parallel run may not be worth the effort. The lighter path is to switch over directly and watch closely through the first week, backed by a rollback plan that is genuinely ready to use.

Safety that nobody ends up needing does not simply disappear. It turns into time the team has to carry during the busiest week of the project. Matching the depth of checking to the real risk of the system is a decision worth making deliberately, rather than going all-in every time out of fear or skipping it out of haste.

What to write into the TOR to get data you can actually use

If the project is still at the scoping stage, three data conditions belong in the TOR (terms of reference) from the first draft. Adding them later, once the vendor has already delivered, is usually hard and comes with a negotiation cost.

First, measurable acceptance criteria for the data. Spell out how correctness will be measured and what the pass threshold is, instead of a broad phrase like "data migrated completely", which is too easy to declare satisfied while the data is still wrong.

Second, a clear statement of who owns data cleansing: the organization or the vendor. Legacy data is usually dirtier than anyone estimated, with duplicate records, values entered in the wrong format, and fields left empty for years. If the TOR does not say who is responsible for cleaning it, the work turns into a mid-project dispute over who should do it and who should pay.

Third, delivery of a mapping document or data dictionary that records which field in the old system goes to which field in the new one, and by what transformation rule. This is what your internal team will rely on once the vendor is no longer around, and it is the only record that explains where today's data came from.

The objection we hear most often

The most common pushback is that the contractor has done many migrations before, so leaving it to them should be enough. It is a fair point, and the answer is not that the contractor cannot do it. The technical side of moving data is something an outside team does well, often better than an internal team that rarely does this kind of work.

What an outside team lacks is the knowledge of which legacy data "looks odd but is right" and which "looks normal but is wrong". A customer with a negative outstanding balance may be perfectly normal for that business, while a record that looks tidy may have been entered incorrectly since last year. That knowledge lives with the internal team that uses the data every day. It is not written down anywhere that could be handed over.

So the job of accepting the data cannot be separated from the job of moving it. The vendor can deliver the goods, but the person who opens the box and signs off that the contents are correct is the business owner of the data. Splitting the work this way takes nothing away from anyone. It simply places responsibility with the people who have enough information to carry it.


In short

The most expensive stretch of a system replacement is usually not the coding or the data migration itself. It is the few hours in which someone has to decide whether you are truly ready to go live. That decision gets much easier when the pass criteria are known in advance and a tested way back is waiting.

If you take one thing from this article, write down what "good enough to pass" means before the migration starts, together with a fallback plan that has been tested and shown to work. Those two things are what keep go-live night from becoming the longest night your team has had.

If your organization is planning a system replacement or drafting the scope of work, and you would like someone to review your migration and cutover plan before it goes out, we would be glad to talk. You can reach us at 088-983-9386 or [email protected], and our office is in Bangkapi, Bangkok. There is no deadline and no need to decide anything quickly. If the conversation shows that it is not yet time to replace the system, we will tell you so.

FAQ: Frequently Asked Questions about This Article

A collection of questions and answers to help you better understand the content of this article.

Tags:

data migrationcutoverdata validationrollback plango-livesystem replacement
Share:

Other Articles

Stay tuned for upcoming articles!