Skip to main content
Customer Story

How a UK Financial Services Firm Migrated 120 Legacy ETL Jobs in 8 Weeks

Gibran Kazi5 October 2026 · 8 min read

In early 2026, a mid-tier UK financial services firm — we will call them "the firm" — had to move roughly 120 SSIS packages off ageing on-prem infrastructure before the end of the year. The packages fed daily risk desks, regulatory reporting and client onboarding. Eight weeks later the estate was fully migrated, every package had passed row-level parity tests before cutover, and the first full regulatory reporting cycle on the new platform ran without incident. This is how it happened.

1. A deadline the business could not move

The firm's on-prem SQL Server stack, which hosted the SSIS estate, had to come off its ageing infrastructure by the end of the year. The packages were not trivial: daily risk feeds feeding the trading desk, regulatory reporting pipelines, and client onboarding workflows that compliance depended on. Moving the platform deadline was not an option, and neither was getting a regulatory submission wrong.

The firm had quotes from two consultancies for a manual rewrite of the estate onto their chosen target, Databricks. Both landed in the same range: 12 to 18 months, at a cost north of £1 million, with no guarantee that the translated logic would match the legacy behaviour — only that it had been tested, somewhere, by someone. For a firm whose regulator expects evidence of correctness, not assurances, that was not a comfortable place to start.

We have anonymised the firm because of commercial sensitivities, but the numbers and the shape of the engagement are real and, for estates of this size, typical. If you want a rough model of what a similar estate costs to migrate, the ETL migration ROI breakdown works through the same economics in detail.

2. The situation: a deadline, a regulator, and no spare capacity

The 120 packages split roughly into three groups. Around half fed daily risk and treasury processes — data landed overnight, transformed, and was consumed by downstream models every morning. A quarter supported regulatory reporting: returns submitted to the FCA and other bodies on fixed schedules, where a silent change in a number is the worst kind of incident. The remainder handled client onboarding and reference data, important operationally but tolerant of a planned cutover window.

The firm's own data team — two data engineers and a platform lead — was fully committed to keeping the current estate running. There was no appetite for a second platform team's worth of headcount, and even less appetite for an 18-month programme that would force them to maintain every business change twice, in SSIS and PySpark, for the duration. Dual maintenance for that long is where manual rewrite programmes quietly bleed money and lose their best engineers.

3. The constraints: audit trail, zero tolerance, freeze windows

Three constraints shaped everything about how this migration had to be run.

Auditability. As an FCA-regulated firm, the firm needed a defensible record of what each legacy package did and how the replacement was shown to do the same thing. "We tested it" is not a control; "we ran both systems on the same inputs and diffed the outputs row by row, and here is the evidence" is. Every parity test run needed to be retained and attributable.

Zero tolerance for reporting errors. The regulatory feeds could not drift, even by a rounding convention. This sounds obvious but it is where most migrations quietly fail: two systems can both be "correct" and still produce different numbers because one truncates and one rounds. The migration process had to surface those differences explicitly, package by package, before cutover — not in the first submission after it.

Freeze windows. The firm operated around month-end, quarter-end and submission calendars. Any cutover activity had to fit into agreed change windows, which effectively meant sequencing the migration by feed criticality and never touching the risk feeds mid-week. An 18-month manual programme spans a lot of freeze windows. An 8-week one spans very few.

4. The approach

The destination was Databricks, with the business logic translated to PySpark notebooks. The firm chose that target before we were involved; the SSIS to Databricks migration approach describes the target-side shape in general terms. What mattered for this engagement was the process, which had five steps.

Step 1: Extract the logic, not just the shell. Each of the 120 packages was parsed to recover the actual transformations — derived columns, lookups, aggregations, conditional splits, slowly changing dimensions — together with the data flow structure and any embedded SQL. Undocumented packages were the norm rather than the exception. Several contained logic nobody on the current team could explain, which is exactly the material that surfaces in month 8 of a manual rewrite and gets priced nowhere.

Step 2: Translate to PySpark. The extracted logic was translated into production-ready PySpark notebooks structured for the firm's Databricks environment, with configuration externalised rather than hard-coded.

Step 3: Generate parity tests with synthetic data. For each package, synthetic test data was generated to model the production data distributions — including the awkward cases: null-heavy columns, boundary values, duplicate keys, unusual character encodings, the lot. The same inputs were run through the legacy SSIS package and the new PySpark code, and the outputs were compared row by row. No production data was used in validation, which kept the whole exercise clean from a data-handling perspective and satisfied the firm's information security review without a fight.

Step 4: Human review and sign-off, per package. The firm's two data engineers reviewed the generated code and, critically, every parity test failure. Failures were almost never "the translation is wrong" in the arithmetic sense; they were places where the legacy behaviour itself was ambiguous or outright buggy — the package that had been silently dropping rows for years, the currency conversion applied twice. Those packages went back with a written recommendation and a human decision. This is the part that matters in a regulated estate: nothing shipped on generated output alone, and every sign-off was attributable to a named engineer.

Step 5: Phased cutover by feed criticality. Onboarding and reference data feeds went first — lower risk, tolerant of a planned window — then the daily risk feeds, then the regulatory reporting pipelines, each after their parity tests passed and inside an agreed change window. Legacy and new pipelines ran in parallel during a short soak period per feed, so rollback at any stage meant turning a switch, not an emergency rewrite.

5. The results

The engagement ran for eight weeks elapsed, from first package extraction to final regulatory feed cutover.

  • Timeline: 8 weeks, against manual rewrite quotes of 12–18 months.
  • People: the firm's two data engineers as reviewers, plus LogicLift. No new hires, no consultancy bench.
  • Parity: 100% of packages passed row-level parity tests before their cutover. There were no "cut it over and see" packages.
  • Post-migration: the first full regulatory reporting cycle on the new platform completed without incident, and there were no reporting incidents attributable to the migration in the following months.
  • Licences and infrastructure: the SSIS estate and its supporting SQL Server infrastructure were decommissioned on schedule, which is where the recurring saving actually comes from.
  • Audit trail: the firm held a complete record — extracted logic, translations, parity test inputs, outputs and sign-offs — per package. Their internal audit reviewed it without escalation.

The honest footnote: two packages were retired rather than migrated, because review showed they produced output nothing consumed. That is a feature of the per-package review step, not an accident — manual rewrite programmes tend to faithfully port dead logic, because nobody is incentivised to look for it.

Facing a migration deadline in a regulated estate?

Book a free 30-minute migration assessment. We will look at your estate size, your regulatory constraints and your target platform, and map the fastest safe route to a parity-tested cutover.

6. What made it work: three lessons

1. The constraint set was treated as requirements, not obstacles. FCA audit trail, zero tolerance for reporting drift, freeze windows — these dictated the shape of the validation and the cutover sequencing from day one. A process that treats regulation as friction will lose to one that treats it as a specification. For any regulatory data migration in the UK, the evidence trail is the deliverable as much as the code is.

2. Parity testing happened before cutover, not after. Running both systems on production-shaped synthetic data and diffing outputs row by row, per package, is what made "100% pass before cutover" a meaningful claim rather than a slogan. It is also what made the auditors happy: the firm could show the evidence, not describe a methodology.

3. Reviewer time was planned in, not hoped for. Two named engineers, with review time protected in their diaries for eight weeks. Human-in-the-loop review fails when it is an afterthought assigned to whoever is least busy; it works when it is a resourced role with authority to reject. The eight-week timeline only held because the firm took that seriously.

7. Could this apply to you?

Honestly, sometimes. This shape of engagement works well when:

  • Your estate is in the 50–300 package range — large enough that a manual rewrite is a programme, small enough that an 8–10 week engagement is realistic.
  • The business logic should be preserved, not redesigned. If you want to re-architect the logic, translate first, redesign deliberately afterwards, on a stable target.
  • You can commit one or two senior engineers to review. That is the real cost on your side, and there is no honest way around it.
  • Your target is reasonably settled. This firm was committed to Databricks. Swapping targets mid-migration is survivable but wasteful.

It is less likely to fit when the estate is a handful of packages (a manual rewrite is simply quicker), or when the organisation wants the migration programme itself to deliver a strategic re-architecture in one step. Those are different jobs, and pretending otherwise is how 18-month programmes happen.

8. Next steps

Gibran Kazi, founder of LogicLift — AI-assisted legacy ETL migration with built-in parity tests.

financial services data migration case studySSIS migration bankregulatory data migration UKETL migration case study