The 10 ETL Migration Mistakes That Blow Up Timelines (and How to Avoid Them)
Most legacy ETL migrations do not fail in a dramatic technical blow-up. They fail for boring, predictable reasons visible from week one: nobody counted the packages, testing amounted to eyeballing row counts, and nobody planned how to go back when the new pipeline proved wrong on day three. Whether your estate is SSIS, Talend, Informatica, DataStage or SAP, and your target is Databricks, Microsoft Fabric, dbt or Snowflake, this list is the shortest guide to ETL migration best practices we know how to write: each mistake, what it is, why it happens, what it costs and the fix.
1. Why this matters now
The pressure to move is real: SSIS is frozen in feature terms, Talend Open Studio users face forced moves to commercial editions, and Informatica PowerCenter customers are being pushed towards IICS with repricing. Databricks, Fabric and Snowflake are where new data engineering talent already lives, and UK salaries have risen enough that "just hire a team to rewrite it" is a different business case than it was three years ago. CFO scrutiny has tightened too: programmes that cannot show a credible timeline, a risk-adjusted cost and evidence of correctness get paused.
2. The status quo, and its flaws
The default approach to a legacy ETL migration is a manual rewrite: engineers open each package, read the logic, and hand-write PySpark, SQL or dbt to reproduce it. For a 100–300 package estate that typically means 12–18 months and a seven-figure all-in cost, with the familiar discovery surprises arriving in month eight rather than month one. Cloud ELT tools help with ingestion but do not translate the business logic locked inside a fifteen-year-old package. Neither path is inherently wrong, but both share the same failure modes: the ten below.
3. The 10 mistakes
10. Rewriting by hand instead of translating the logic
The mistake: treating a migration as a green-field rewrite, with engineers hand-transcribing each package into the new technology.
Why it happens: "we will rebuild it properly" is easy to say and easy to fund. Nobody has to explain translation tooling to a sceptical steering committee.
What it costs: at the usual rule of thumb of one complex package per engineer per one to two weeks, a 200-package estate is a 12–18 month programme in which every business change must be made twice. Assisted translation of estates this size typically runs in weeks.
The fix: extract and translate the existing logic wherever it should be preserved; reserve engineering effort for the packages that genuinely need redesign. Rewrite is a design decision, not a default.
9. Migrating data but not the business logic locked in packages
The mistake: moving the data successfully while leaving the transformation logic behind, or assuming an ELT tool's ingestion connectors constitute a migration.
Why it happens: data movement is visible and demonstrable; business logic is invisible until something downstream breaks. Reporting "80% of feeds migrated" is easy while the hard 20% — derived columns, lookups, slowly changing dimensions — remains untouched.
What it costs: the new platform produces numbers that do not match the old one; finance spots a discrepancy, confidence collapses, and the programme spends months forensically comparing outputs.
The fix: inventory the logic, not just the data flows. Every derived column, aggregation and conditional branch in a legacy package is a requirement, whether or not anyone documented it.
8. No inventory — not knowing what you actually have
The mistake: budgeting and planning against a wiki page last updated in 2019, when the real estate includes orphaned packages, undocumented jobs and schedules nobody owns.
Why it happens: estates grow organically for a decade or more; jobs get cloned and modified without being reconciled, and nobody owned keeping the inventory current.
What it costs: scope discovered mid-programme is the most common cause of overrun: the undocumented nightly job feeding a regulatory report, surfacing in month seven, is a slip measured in months, not weeks.
The fix: start with a mechanical discovery pass over the estate itself — enumerate every package, job and schedule, and classify each as active, dormant or orphaned. Expect roughly a third more logic than the documentation suggests.
7. Skipping parity testing — or testing only row counts
The mistake: declaring the migration correct because the new pipeline loads "about the same number of rows".
Why it happens: row counts are cheap to check, easy to report and reassuring to non-technical stakeholders. Row-level correctness is harder.
What it costs: row counts say nothing about values. A subtly wrong join, a rounding difference or a lost NULL can leave row counts identical while every figure in the warehouse is wrong — and the defect only surfaces when a business user notices a number that looks odd.
The fix: parity testing means running the same inputs through the legacy pipeline and the new pipeline and diffing the outputs row by row, column by column, with tolerances made explicit and justified. Anything less is a hypothesis, not a test.
6. Big-bang cutover instead of phased parallel running
The mistake: switching everything off on a Friday and switching the new platform on the following Monday.
Why it happens: dual running costs money and attention, so programmes under budget pressure cut it — helped by the psychological pull of a clean, definitive cutover date.
What it costs: when (not if) something differs in week one, there is no baseline to compare against and the rollback is a re-migration under time pressure. Big-bang cutovers are how "six more weeks of parallel running" becomes "nine months of incident management".
The fix: migrate domain by domain, run old and new in parallel for a defined soak period, and cut over each domain independently once parity is evidenced on production-shaped data. Boring, incremental and almost impossible to fail catastrophically.
5. Ignoring rollback planning
The mistake: writing the go-live runbook but never the go-back runbook.
Why it happens: rollback planning feels like planning for failure, and programme cultures reward optimism — it also forces awkward questions about data reconciliation that everyone would rather defer.
What it costs: the first serious production defect arrives with no tested route back. The team improvises — replaying extracts, reconciling partially-loaded targets — and recovery takes days. Stakeholder confidence takes longer to recover than the data.
The fix: define rollback before cutover, in writing, per domain: what triggers it, who decides, how data is reconciled, how the legacy pipeline restarts. Then rehearse it — a rollback plan that has never been executed is a hypothesis.
4. Under-resourcing the human review
The mistake: treating human-in-the-loop review as an afterthought — "we'll have someone check the output at the end".
Why it happens: once tooling "does the work", it is tempting to count the review as zero-cost. Nobody books reviewer time, and nobody senior is named as accountable for sign-off.
What it costs: unreviewed generated code inherits the defects of unreviewed hand-written code, with extra confidence. Edge cases — NULL handling, locale-specific dates, that one weird package from 2014 — pass silently, and the eventual review is rushed and done by whoever is free.
The fix: plan review capacity as a first-class resource from day one. Name a senior engineer as sign-off owner per domain, protect their time, and make sign-off a gate in the plan rather than a checkbox at the end. It is also where genuinely dead legacy logic should be caught and retired, not translated faithfully into the new platform.
3. Underestimating regulatory and audit sign-off lead times
The mistake: assuming the data team's definition of "done" is the organisation's definition of done.
Why it happens: in regulated sectors a migration is not complete when the pipeline runs; it is complete when compliance, risk and internal audit have seen evidence of correctness and signed off — and data teams do not control those calendars.
What it costs: the cutover lands on schedule and the programme then waits two to three months for sign-off while both estates must be maintained. Discovering this in month ten rather than month two routinely costs a quarter.
The fix: engage compliance and audit at the start, agree what evidence they need, and generate it as a by-product of the migration. Row-level parity evidence, produced continuously, is far easier to sign off than a verbal assurance that "we checked it carefully".
2. The training gap — new platform, team that only knows the old tool
The mistake: migrating the estate but not the people, then discovering the support team cannot operate the new platform.
Why it happens: programmes are scoped around technology deliverables; training is a soft, easy-to-defer line item, and hiring is assumed to close the gap.
What it costs: after cutover every incident routes to the two people who understand Spark, the legacy experts' knowledge evaporates as they are reassigned, and shadow pipelines appear as teams work around the platform team.
The fix: pair legacy experts with platform engineers throughout, treat the translated code as a training artefact, and run supported hypercare after cutover — the migration is a knowledge-transfer exercise whether you plan it as one or not.
1. No total-cost modelling
The mistake: costing the migration as engineering effort plus new platform licences, and stopping there.
Why it happens: the hidden costs — dual running, parallel maintenance of two estates, regression testing effort, compliance sign-off delay, attrition — are spread across different budgets, so nobody sees them in one place.
What it costs: the classic data migration pitfall is a programme that hits its engineering budget with no end in sight, because the real cost was twice what anyone modelled. Boards lose trust, funding gets re-examined, and the "on budget" migration becomes the reason the next one is never approved.
The fix: model total cost before you start: legacy run-cost until decommission, dual running, parallel maintenance, testing and compliance overhead, and the new platform's run-cost after. The worked example in our ETL migration ROI breakdown is built for exactly this exercise — do it before any business case is written, whichever delivery route you choose.
4. The LogicLift approach
Everything above is tool-agnostic advice. It is fair to say, though, how LogicLift's assisted-migration methodology is designed around these specific failure modes, since that is the point of the approach.
LogicLift reads legacy packages — SSIS, Talend, Informatica, DataStage, SAP — and extracts the business logic itself: derived columns, lookups, aggregations, SCD handling. That addresses mistake 10 (translation rather than hand-rewrite), mistake 9 (the logic moves with the data) and mistake 8 (the estate inventory is a mechanical first output, so orphaned packages surface in week one).
The extracted logic is translated into target-native code — PySpark notebooks for Databricks, notebooks and pipelines for Fabric, dbt models where a warehouse is the target — and parity tests are generated with synthetic data, diffing legacy output against new output row by row. That is mistake 7 designed out of the process, and the continuous parity evidence is the same evidence that shortens mistake 3's audit sign-off. Phased parallel running with per-domain rollback plans is expected, and human-in-the-loop review is a scheduled, resourced gate with a named sign-off owner — mistakes 6, 5 and 4 treated as first-class concerns, not risks logged and forgotten.
For estates in the 100–300 package range, a typical engagement runs 6–10 weeks to production-ready, parity-tested code, versus 12–18 months for a manual rewrite. The two mistakes this does not remove are yours: total-cost modelling and training are organisational. The honest ROI summary is in our ETL migration ROI post; the SSIS to Databricks use-case page shows the process end to end.
Planning a legacy ETL migration?
Book a free 30-minute migration assessment. We will map your estate, flag which of these ten mistakes your current plan is most exposed to, and outline the fastest safe route to a parity-tested cutover.
5. ROI summary
- Every mistake on this list has the same financial signature: elapsed time. A 12–18 month manual rewrite carries 12–18 months of dual maintenance, licence fees on the retired platform, attrition risk and delayed benefit.
- Methodologies that avoid these mistakes compress that to weeks — an illustrative estate of 100–300 packages typically sees payback inside the first year rather than beyond the third.
- The calculation that matters is not the engineers' day rate but the risk-adjusted total cost, the time to cutover, and the evidence you can show a regulator, an auditor or a board. Get those three right and most failure modes are already closed off.
6. Further reading and next steps
- See how LogicLift migrates SSIS to Databricks
- Model your migration cost: the ETL migration ROI breakdown
- Work through the modernisation checklist before you migrate
- Book a migration assessment
- LogicLift home page
Gibran Kazi, founder of LogicLift — AI-assisted legacy ETL migration with built-in parity tests.