← Blog

In-House vs Outsourced Data Migration: Choosing the Delivery Model

Kaan Dincer
Founder & CEO, Settle. Previously ran Fortune 500 data migrations at Deloitte.
August 5, 2026

The in-house versus outsourced question for data migration is usually decided by default rather than on its merits: the work arrives bundled inside the system integrator's statement of work, and nobody revisits it. There are actually four delivery models, not two, and the right one turns on three variables: how many source systems you have, how hard your deadline is, and whether your subject matter experts can staff four to eight load-and-fix cycles while running the business.

Key takeaway: The data workstream deserves its own sourcing decision, separate from choosing your implementation partner. Ownership of what the data means stays with your business under every model. What varies is who executes the movement and who proves the result. The most common failure is not picking the wrong model. It is never noticing there was a choice.

The decision nobody makes deliberately

When an organization signs an ERP, CRM, or HRIS implementation, the data migration line is already inside the integrator's proposal. Procurement prefers one contract. Leadership prefers one accountable party. The integrator prefers to keep the scope. All three preferences are reasonable, and together they produce a sourcing decision that was never examined.

A delivery model for the data workstream defines who executes the extraction, mapping, transformation, loading, and validation of your data, and who is accountable for proving the result at cutover. It is a separate decision from which implementation partner you hire and which system you buy.

The bundled default is sometimes the right answer. The problem is structural: because the data line is one item inside a larger bid, it is priced to help win the implementation, and the risk that cannot be estimated at bid time, the cleansing effort and the reload cycles, tends to move into change orders. The cost build-up behind that pattern is worth understanding before you accept the default, because it explains why the cheapest data line at signing is frequently the most expensive one at go-live.

The four ways to source the data workstream

Bundled with the system integrator

What you are buying is coordination: one contract, one program plan, one party accountable for the whole implementation. That is genuinely valuable, and for a single-source migration with standard objects and an SI whose data practice answers the fifteen vendor questions precisely, it can be the right call.

Where the risk sits is in priorities. Inside an implementation, the data workstream competes for the integrator's strongest people with the configuration workstream, and configuration usually wins, because configuration is what the demo runs on. The data line gets the remaining attention, and the parts of it that resist estimation get absorbed later as change orders.

If you stay bundled, make the data line separable inside the SOW: its own acceptance criteria, its own stated cycle count with a price for additional cycles, and the reconciliation evidence package as a named deliverable. Bundled with structure is a legitimate model. Bundled by default is a hope.

A specialist alongside the integrator

What you are buying is a party whose only deliverable is the data. The specialist runs profiling, mapping, transformation, and validation with its own acceptance criteria, while the integrator owns configuration and the overall program. The data workstream stops competing for attention because it has its own owner.

Where the risk sits is in the seams. The interface between specialist and integrator has to be defined explicitly: who loads into which environment, whose staging is authoritative, where the defect queue lives, and who calls a record closed. Undefined seams turn into finger-pointing at exactly the moment the schedule is tightest.

This model fits multi-source consolidations, hard external deadlines, audit-heavy environments, and any situation where the integrator's answers to the vendor questions were vague. It is also the model most integrators will quietly accept once a client asks for it, because it moves the least estimable risk off their book too.

Fully in-house

What you are buying is retained knowledge and no vendor margin. Your team learns the target system's data model deeply, the transformation logic stays in the building, and every rule discovered is a rule you keep.

Where the risk sits is capacity, not skill. The people who know what your data means are the people running the business, and a migration needs them for interviews, adjudication, and testing across every load cycle. First-time teams also learn the reload loop by living it, which means the four-to-eight-cycle pattern arrives as a surprise instead of a plan.

This model fits a single clean source, a soft deadline, and a team that can genuinely calendar the hours. It also fits organizations for whom migration is recurring rather than one-time, where building the muscle pays back on every subsequent project, and where the build-versus-buy question becomes about tooling rather than staffing.

In-house execution with independent validation

What you are buying is the smallest external footprint that still produces proof. Your team does the mapping and the loads. An outside party runs the profiling, operates the validation gates, and produces the reconciliation evidence, at full-dataset grain rather than samples, using the discipline described in how to validate data after a migration.

Where the risk sits is authority. If the validating party cannot stop a load, the validation is decoration. The gate needs teeth in the project charter: defined pass criteria, and a rule that a failed gate halts the schedule rather than annotating it.

This hybrid fits capable internal teams facing a deadline or an audit requirement that makes self-certification uncomfortable. It is the cheapest way to add rigor without handing over execution.

Ownership does not change with the model

Whichever model you choose, one thing cannot be outsourced: deciding what the data means.

The data workstream owner split holds under every delivery model: the business owns the meaning of the data, including field definitions, survivorship rules, and the archive boundary, while the executing party, internal or external, owns the movement and the proof.

Name a single accountable business owner per object and put the names in the plan: supply chain or engineering for the item master, sales operations for customers, the controller for the chart of accounts and balances. Vendors can propose survivorship rules. They cannot sign off on them, because when two source records conflict, the right answer is a business judgment about which relationship, which price, which history is true. Programs that outsource that judgment do not actually outsource it. They just have it made implicitly by whoever wrote the transformation code. The fuller version of this argument is in the ERP data migration pillar.

Can we do it ourselves? The honest test is arithmetic

The in-house question is almost never about skill. Competent data engineers exist in most mid-market IT teams, and the target vendors document their load interfaces. The question is whether the organization can staff the iteration.

A typical mid-market program pulls two to five subject matter experts at twenty to forty percent allocation for six to twelve months, across four to eight full cycles of load, triage, correction, and reload. Those SMEs are, by definition, your best operators, because they are the people who know what the data means. The test is a calendar, not a confidence level: block those hours across those months for those named people, and see whether the business still runs. If the calendar does not close, the answer is no, regardless of how capable the team is. If it closes, in-house is a real option, and the remaining question is whether you want to learn the reload loop on a live project or compress it with tooling and outside validation.

One more distinction that changes the answer: frequency. If this migration is a one-time event, an in-house build creates capability you will use once. If migrations recur, because you onboard acquired companies or move customer data as part of your operating rhythm, building the internal muscle is an investment rather than a detour, and the sourcing question shifts from people to platform.

Services versus consulting firms: hours or outcome

Within the outsourced options, there is a distinction that matters more than brand names.

A consulting firm sells staffed hours: skilled people, billed by time, accountable for effort and advice. A data migration service sells a delivered outcome: a validated dataset with reconciliation evidence, priced against the result rather than the hours it takes.

Neither is dishonest. They are different products, and the buying mistake is purchasing one while expecting the other. If your team needs augmentation, judgment on tap, and flexibility about where the engagement goes, you want consulting, and you should expect to manage the work. If you need a proven dataset in the target by a date, with evidence, you want a service, and the vendor should carry the iteration risk inside a fixed price.

The tell is in the commercial structure. Ask what happens to the price if the project needs two more load cycles. A consulting answer is more hours. A service answer is nothing, because the cycles are the vendor's problem. Settle is built as the second kind: the data workstream delivered, not staffed, fixed and scoped, at roughly half the equivalent consulting-led cost, with the validation gates and the evidence package as the deliverable.

Decision framework: matching the model to the situation

If you have a single source, clean master data, and a soft deadline, run it in-house, and spend your external budget, if any, on a short independent validation of the result.

If you are consolidating multiple sources, or migrating as part of M&A, bring in a specialist for the data workstream. Survivorship across sources is the hard problem, it needs a dedicated owner, and it is the category where bundled data lines most reliably overrun.

If you face a hard external deadline, an end of support date, a TSA exit, a lease expiry, do not learn on the job. Choose a specialist or a structurally separated bundled line, and buy the reload cycles explicitly, because cycles are what you will run out of.

If you operate under audit or regulatory requirements, choose at minimum the hybrid model. Self-certified migration evidence is weak evidence, and the reconciliation package needs to be produced by a party with standing to fail the load.

If migration recurs in your operating model, invest in the in-house muscle and evaluate platforms rather than renting consultants per event. The economics of the two paths diverge more with every migration you run.

If you stay bundled with your integrator, and often you reasonably will, separate the data line inside the SOW: acceptance criteria, cycle count, cleansing responsibility with an assumed defect volume, and the evidence package as a named deliverable. How that line usually gets written, and why it is so often the most underquoted item in the SOW, is its own subject, and we will treat it separately.

The four models at a glance

ModelWhat you are buyingWhere the risk sitsWhen it fits
Bundled with the SIOne contract, coordinated programData competes with configuration for attention; unestimable work drifts to change ordersSingle source, standard objects, strong SI data practice, structured SOW
Specialist alongside the SIA party whose only deliverable is the dataSeams with the integrator must be defined explicitlyMulti-source, M&A, hard deadlines, audit requirements
Fully in-houseRetained knowledge, no vendor marginCapacity of SMEs; learning the reload loop liveSingle clean source, soft deadline, recurring migrations
In-house plus independent validationRigor and evidence without handing over executionValidator must have authority to stop a loadCapable team plus a deadline or audit pressure

What the default looks like from inside

At Deloitte, the data workstream came to us bundled by default, inside programs sold on configuration and process design. The clients who did best were the ones who treated the bundling as a choice they were actively making: they set separate acceptance criteria for the data line, asked how many cycles the estimate carried, and named their own business owners per object before we mapped anything. The clients who did worst never asked, and discovered the shape of the data work at the same time we did, one load at a time, on their schedule and their invoice.

The model itself mattered less than whether it had been chosen. Every one of the four can succeed. The unexamined default is the only one that fails quietly.

Frequently Asked Questions

Should data migration be done in-house or outsourced?

It depends on capacity more than skill. In-house works when you have a single clean source, a soft deadline, and subject matter experts who can genuinely calendar twenty to forty percent of their time across six to twelve months and four to eight load cycles. Outsource when sources multiply, deadlines harden, or the SME calendar does not close. A hybrid, in-house execution with independent validation, covers many cases in between.

Can we do the data migration ourselves?

Often yes on skill, and the honest test is a calendar. Block the interview, adjudication, and testing hours for your named subject matter experts across every planned load cycle, and check whether the business still runs. If the calendar closes, in-house is real. If it does not, the constraint is capacity, and no amount of engineering talent substitutes for the people who know what the data means.

Who owns data migration in an ERP project?

Split the ownership. The business owns what the data means: field definitions, survivorship rules, and the archive boundary, with a single named owner per major object. IT, or the external party, owns the movement and the proof. This split holds under every delivery model, and the meaning half cannot be outsourced, because conflicts between source records are business judgments, not technical ones.

What is the difference between data migration services and consulting firms?

Consulting firms sell staffed hours: skilled people billed by time, accountable for effort. Migration services sell a delivered outcome: a validated dataset with reconciliation evidence at a fixed price, with the vendor carrying the iteration risk. The tell is what happens to the price when the project needs two more load cycles. Hours means consulting. Nothing means a service.

How do we choose the data migration team for a system transformation?

Decide the delivery model first, then the party. The four models are bundled with the integrator, a specialist alongside, fully in-house, and in-house with independent validation. Choose using three variables: source system count, deadline hardness, and SME capacity. Then evaluate whichever parties fit the model with the same question set, starting with how many load cycles their estimate assumes.

Should the data workstream be inside the system integrator's SOW?

It can be, and often reasonably is, provided it is separable within the document: its own acceptance criteria, a stated cycle count with pricing for additional cycles, cleansing responsibility with an assumed defect volume, and the reconciliation evidence package as a named deliverable. A data line with those elements is a chosen model. A data line without them is an unexamined default.

Bringing it together

In-house versus outsourced is the wrong binary. The real decision has four options, three deciding variables, and one constant: the business owns what the data means under every model, and the sourcing choice only determines who moves it and who proves it. Make the choice on purpose. The default is the only model with an unbounded failure mode, because nobody is watching the line that nobody chose.

If you are deciding how to source the data workstream on an upcoming transformation, Settle will scope it against your actual sources and constraints, delivered, not staffed. Book a scoping call.