Questions to Ask a Data Migration Vendor Before You Sign
The single most important question to ask a data migration vendor is how many load-and-fix cycles the quote assumes, because cycle count is where migration budgets and schedules actually diverge. The fourteen questions that follow it exist for the same reason: migration proposals are priced on the estimable work and delivered on the unestimable work, and the questions below force the second category into the open before you sign rather than after.
Key takeaway: These questions work on any vendor quoting the data line. That might be your system integrator's data team, a consulting firm, or a specialist platform. In most implementations the data workstream sits inside the SI's statement of work, which makes these the questions you ask your implementation partner about one line of their SOW, not just questions for a standalone shop. Good vendors answer them precisely. The evasions are just as informative as the answers.
Who counts as a data migration vendor
Very few organizations hire a standalone data migration company. The data workstream usually arrives bundled inside a larger engagement, which is exactly why it gets underexamined.
A data migration vendor is whoever is contractually accountable for extracting, mapping, transforming, loading, validating, and reconciling your data into the target system. In practice that is one of three parties: the data practice inside your system integrator, an independent consulting firm, or a specialist migration provider.
Ask these questions of whichever party is quoting the work. If the answer is "the SI handles that," ask the SI. The point is not which category of vendor you chose. The point is that someone has priced this work, and the price rests on assumptions that are almost never written down.
The estimate
1. How many load-and-fix cycles does this quote assume, and what does an additional cycle cost?
Implementations commonly run four to eight rounds of load, error triage, correction, and reload, and each round carries real hours. This is the largest driver of cost variance on the data line. A strong answer is a number, a description of what one cycle includes, and a stated price or mechanism for cycles beyond it. A weak answer is "as many as it takes" with no commercial consequence, which means the risk is unpriced, or "we rarely need more than two," which means the estimate is optimistic.
2. Can you show me the estimate build-up: objects, hours per object, and rate?
Every migration quote is rate times hours underneath, with hours driven by object count and cycle count. A strong answer walks you through the arithmetic: how many objects, hours per object for mapping and transformation, hours per test cycle, blended rate, and the assumptions each rests on. A weak answer is a single number defended as experience with similar projects. A number you cannot decompose is a number you cannot compare against any other bid.
3. Who fixes bad source records, and what defect volume does the estimate assume?
Most SOWs assign source data cleansing to the client, which sounds reasonable until profiling returns forty thousand exceptions and your team owns all of them. A strong answer includes a responsibility matrix naming who corrects what, an assumed defect volume, and what happens commercially if reality exceeds it. A weak answer is "data quality is the client's responsibility" with no volume assumption attached, which converts your staff into an unbudgeted line item.
Discovery
4. Do you profile full production volumes during discovery, or work from samples and documentation?
The exceptions that stall migrations are rare by definition, and rare defects are structurally invisible to samples. A strong answer is that full extracts get profiled during discovery, with findings quantified by exception class before mapping begins. A weak answer is a documentation review plus a sample walkthrough with your subject matter experts. Documentation describes how the system was designed. Profiling reveals how it was used.
5. When in the schedule do we first load full production volume into staging?
This one question predicts the project's shape better than any other. A strong answer is a date inside the first quarter of the timeline. A weak answer is that full loads come after mapping sign-off, somewhere in the back half. Every week between project start and the first full load is a week in which the hardest ten percent of your data remains undiscovered, and the late discovery of exceptions is the most common failure pattern in this work.
6. How do you capture business rules that exist only in people's heads?
Every legacy system carries rules nobody documented: a status field overloaded to mean something else, a free-text field holding structured data, a code that means one thing in the manual and another in practice. A strong answer names a method, usually operator interviews combined with anomaly profiling, and a deliverable, usually a rule inventory with owners. A weak answer is "your team provides the business rules," which assumes the knowledge is written down. If it were written down, you would need less help.
Validation
7. At what grain do you validate: totals, samples, or every record?
Totals reconcile while individual records are wrong, because errors of opposite sign net out. A strong answer is record-level comparison across the full dataset, typically by normalizing and hashing both sides, alongside control totals and count reconciliation. A weak answer is "we reconcile totals and spot-check records." On the size of sample needed to catch rare defects, the arithmetic is unforgiving, and we walk through it in how to validate data after a migration.
8. Is the reconciliation evidence package a named deliverable?
At cutover, someone must be able to prove the migration was complete: count equations per object, control totals, the exception log with resolutions, and sign-offs. Eventually an auditor will ask. A strong answer is yes, as a named deliverable with a defined format, produced at cutover. A weak answer is "we can provide reports on request," which means the evidence gets reconstructed later at your expense, if it can be reconstructed at all.
9. When a record fails validation and gets corrected, what happens next?
The correct answer is that the corrected record re-enters the full validation gate and earns a pass like any other record. Fixed is not passed. A corrected record has failed once already, and corrections made under schedule pressure are how one defect becomes three. A weak answer is a puzzled look, or "once it's fixed it's done." This question sounds pedantic and is not: how a vendor treats corrected records tells you whether their validation is a gate or a formality.
10. Do you simulate real business workflows against migrated data before go-live?
A record can satisfy every structural constraint and still be unusable: an open purchase order that cannot be received against, a customer that cannot be invoiced. A strong answer is a defined transaction set executed in staging before user acceptance testing, covering receiving, shipping, invoicing, and a period close where relevant. A weak answer is "that is covered in UAT," which means your users find the failures, on the clock, at the end.
Delivery
11. Who exactly does the work, and are they named in the SOW?
The team that pitched is not always the team that delivers. A strong answer names individuals in the SOW with their allocation percentages and prior programs. A weak answer is "resources will be assigned at project start." On a workstream where most of the value is judgment about ambiguous data, the specific people are the product.
12. What is the rollback plan if cutover validation fails?
A strong answer includes documented go and no-go criteria, a tested rollback procedure, and a named decision owner for cutover night. A weak answer is "we have never had to roll back," which is either untrue or evidence of go-lives that should have been stopped. The existence of a rehearsed rollback is one of the better proxies for overall delivery discipline, and it belongs in the cutover plan from the start.
13. What artifacts do we keep when the engagement ends?
Mappings, transformation rules, exception logs, validation results, and the evidence package should remain with you in usable form, because you will migrate again and because audits outlive engagements. A strong answer is a defined handover set in open formats. A weak answer is tooling so proprietary that the logic leaves when the vendor does, which converts a one-time engagement into a dependency.
Commercial structure
14. Why is this priced fixed, or why is it priced time-and-materials?
The structure should match how well the scope is known. Fixed price on a well-defined scope buys schedule certainty at a premium. Fixed price on an undefined scope buys the vendor's worst-case assumption. Time-and-materials with no stated cycle assumptions is an open meter. A strong answer justifies the structure against scope certainty, often discovery on T&M converting to fixed price once object and exception counts are known, with change triggers defined in writing.
15. What do you need from our team, in hours per week, by role?
The vendor's fee is only part of the true cost. Your subject matter experts will spend real hours in interviews, adjudication, and testing, and they are the same people running the business. A strong answer is a staffing expectation by role and phase, in hours per week, that you can put on a calendar. A weak answer is "minimal involvement from your side." No honest migration needs minimal involvement from the people who know what the data means.
The fifteen questions at a glance
| # | Question | A strong answer includes |
|---|---|---|
| 1 | How many load-and-fix cycles does the quote assume? | A number, cycle contents, price per additional cycle |
| 2 | Can you show the estimate build-up? | Objects, hours per object, cycles, blended rate |
| 3 | Who fixes bad source records? | Responsibility matrix with assumed defect volume |
| 4 | Do you profile full volumes in discovery? | Full-extract profiling, findings by exception class |
| 5 | When is the first full-volume load? | A date in the first quarter of the schedule |
| 6 | How do you capture undocumented rules? | Operator interviews plus profiling, a rule inventory |
| 7 | At what grain do you validate? | Record-level, full dataset, plus control totals |
| 8 | Is the evidence package a named deliverable? | Yes, defined format, produced at cutover |
| 9 | What happens to corrected records? | Full revalidation; fixed is not passed |
| 10 | Do you simulate workflows before go-live? | Defined transaction set run in staging, pre-UAT |
| 11 | Is the delivery team named in the SOW? | Named individuals with allocations |
| 12 | What is the rollback plan? | Written criteria, tested procedure, decision owner |
| 13 | What artifacts do we keep? | Mappings, rules, logs, evidence, open formats |
| 14 | Why this commercial structure? | Structure matched to scope certainty, change triggers |
| 15 | What do you need from our team? | Hours per week, by role, by phase |
How to read the answers
You are not scoring for perfection. You are scoring for precision. A vendor who answers with numbers, dates, and named deliverables has done this work enough times to know where it goes wrong, and has priced that knowledge in. A vendor who answers in reassurances has priced the mapping and is planning to absorb the rest as change orders, which means you are the one absorbing it.
Notice also which questions produce discomfort. Questions one, five, and seven are the ones that most reliably separate bids, because they are the three places where the standard delivery model hides its risk: unpriced cycles, late discovery, and thin validation.
At Deloitte I sat on the other side of this table, and the clients who asked question five were the ones we scoped most carefully, because they had told us, in one question, that they knew where the bodies were buried. The clients who asked only about rates got estimates that were accurate about rates.
Frequently Asked Questions
What questions should I ask a data migration vendor?
Start with the assumptions under the price: how many load-and-fix cycles the quote includes, what the estimate build-up looks like, and who fixes defective source records. Then probe the method: whether full volumes are profiled in discovery, when the first full load lands in the schedule, and at what grain validation runs. Finish with delivery and commercial structure: named team, rollback plan, retained artifacts, and why the pricing structure fits the scope.
Should we use our system integrator for data migration or a specialist?
Ask both the same fifteen questions and compare the precision of the answers rather than the category of the vendor. Many SI data practices answer them well. The structural risk with a bundled data line is that it is priced to win the overall implementation, which pushes cycle and cleansing risk into change orders. The structural risk with a specialist is coordination overhead with the SI. Whichever you choose, the data workstream needs its own owner, its own acceptance criteria, and its own validation gate.
How do I evaluate data migration services?
Evaluate the assumptions, not the credentials. Certifications and logos tell you a firm has sold data migration services before; the cycle count, the profiling method, and the validation grain tell you how they will handle yours. Ask for the estimate build-up and check the arithmetic against the ranges in a published cost model. A bid meaningfully below the others is usually missing cycles or pushing cleansing onto your staff, not discovering efficiency.
What are red flags in a data migration proposal?
A single-number estimate with no build-up. No stated cycle assumption. Cleansing assigned to the client with no volume estimate. Full-volume loads scheduled after mapping sign-off. Validation described as reconciliation plus spot checks. No named team. No rollback procedure. Any guarantee of accuracy percentages, since honest vendors describe method and evidence rather than promising outcomes.
How many mock loads should a vendor include?
Most programs need between four and eight full load-and-fix cycles under a conventional approach, so a quote assuming two is optimistic unless the vendor validates the full dataset against target rules before loading, which is what removes cycles. The number itself matters less than whether it is stated, priced, and tied to a mechanism for exceeding it.
What should the data migration section of an implementation SOW contain?
The object list, source list, and history window. The assumed cycle count and the price of additional cycles. The responsibility matrix for cleansing with an assumed defect volume. The validation method and grain. The reconciliation evidence package as a named deliverable. Named delivery staff. Rollback criteria and procedure. Client staffing expectations by role. If those are present, most of the fifteen questions are already answered in writing, which is where you want them.
Bringing it together
Every question on this list exists because of the same underlying fact: migration work is priced on the ninety percent that is estimable and lost on the ten percent that is not. The questions force that ten percent into the quote, where it can be compared, priced, and owned, instead of into month seven, where it can only be absorbed.
For what it is worth, this is also the standard we hold ourselves to. Settle validates the full dataset against the target system's rules before load rather than discovering them from load failures, which is what takes cycles out of the schedule, and engagements are fixed and scoped, at roughly half the equivalent consulting-led cost. Ask us all fifteen. Book a scoping call.