Databricks vs Fabric vs Snowflake for ETL Workloads: An Honest 2026 Comparison
If you are migrating a legacy ETL estate — SSIS, Talend, Informatica, DataStage — your shortlist almost certainly contains the same three platforms: Databricks, Microsoft Fabric and Snowflake. All three are credible. All three will run your workloads. And all three will happily take your money for the next five to ten years. This post is an ETL workload platform comparison with one specific lens: not greenfield analytics, but the data engineer porting fifteen-year-old packages whose business logic must survive the journey intact.
1. The question everyone migrating legacy ETL has to answer
The problem is that most "Databricks vs Fabric vs Snowflake" comparisons are written for greenfield analytics teams choosing a lakehouse, not for data engineers porting fifteen-year-old SSIS packages. A generic platform comparison tells you very little about what your migration will actually look like: what your translated pipelines will be written in, who will be able to maintain them, what the compute will cost at 6 a.m. on a Monday when the overnight batch runs, and what happens when the platform you picked turns out to be the wrong one.
This post assumes you are not designing a modern data stack from a blank page; you are moving a working (if creaking) estate, with real business logic that must survive the journey intact.
2. Why this choice matters more than most platform decisions
Platform choices made during a migration have a longer half-life than almost any other architectural decision, for three reasons.
- The target determines the shape of the translated code. Business logic ported to PySpark notebooks does not port neatly to T-SQL stored procedures, and a semantic model built in Fabric does not move to Snowflake without a rebuild. You are not just choosing infrastructure; you are choosing the language your pipelines will be written in for the next decade.
- Migration is when skills get locked in. If your translated estate lands in Spark and your team is a Microsoft SQL shop, you have created a permanent dependency on contractors or a permanent training programme. The reverse is also true.
- The economics harden quickly. Once you have moved 200 pipelines, re-platforming is a second migration, and nobody funds that. The time to be honest about "Databricks vs Snowflake cost" or Fabric capacity economics is before cutover, not after.
The rest of this post covers the differences that actually matter for migrated ETL workloads, then a decision matrix by estate type and team skills.
3. The real differences for migrated ETL workloads
Compute model and pricing
The three platforms price compute in fundamentally different units, which is why list-price comparisons mislead.
- Databricks charges in Databricks Units (DBUs), a consumption metric on top of cloud VM or serverless compute. The Photon engine accelerates SQL and Spark workloads considerably, but you pay for it through DBU multipliers. Costs are workload-shaped: a heavy transformation hour costs what it costs, and idle time is (mostly) not billed. This is flexible and, for spiky batch estates, often the cheapest honest answer — but forecasting requires modelling actual pipeline behaviour, not reading a pricing page.
- Microsoft Fabric sells capacity: you buy an F SKU (measured in capacity units) and everything — Spark notebooks, T-SQL warehouses, pipelines, Power BI — draws from the same pool. The "one lake, one copy" pitch is genuine: OneLake stores data once and every engine reads it, which removes the duplication and egress costs that accumulate in multi-engine estates. The trade-off is that a fixed capacity pool is either under-utilised (wasted budget) or over-subscribed (workloads queue and slow down at peak). Right-sizing capacity for a migrated batch estate takes a few months of observation.
- Snowflake charges in credits consumed by virtual warehouses, billed per second while running (with a per-warehouse minimum). Compute and storage are cleanly separated, auto-suspend is mature, and costs are predictable for SQL-shaped workloads. Snowpark brings Python, Java and Scala execution inside the warehouse, at credit prices that reward efficient code and punish naive porting of row-by-row SSIS logic.
The honest summary: no platform wins on sticker price. Each wins for a different workload shape, which is why section 5 argues for modelling cost per workload rather than comparing rate cards.
Transformation paradigm
This is the difference that shapes your migration more than any other.
- Databricks is a Spark-first platform. Migrated business logic — derived columns, lookups, aggregations, slowly changing dimensions — lands as PySpark or Spark SQL in notebooks. If your SSIS or Talend estate contains heavy data-flow logic, Spark is a natural home for it, and the port is usually more faithful than a SQL rewrite.
- Fabric offers two paths: Spark notebooks (similar to Databricks) and a T-SQL/warehouse path, where transformations live in stored procedures, views and semantic models consumed directly by Power BI. For estates whose logic is mostly SQL-flavoured — a large share of SSIS estates, honestly — the T-SQL path is shorter and more maintainable. The semantic model layer is a genuine differentiator if BI consumption is the point of the estate.
- Snowflake is warehouse-first. SQL is the native language; Snowpark and dbt models cover what pure SQL cannot, and dbt in particular has become the default transformation layer for Snowflake estates. Well-suited to estates being modernised into clean, testable, SQL-centric pipelines — less suited to estates with heavy procedural or streaming logic.
Orchestration
- Databricks Workflows handles notebook and job dependencies natively, with reasonable scheduling and monitoring. Complex inter-job dependencies may still push you toward external orchestration.
- Fabric pipelines (the Data Factory heritage) are the strongest of the three for visual, dependency-heavy orchestration — familiar territory for teams coming from SSIS control flows or Talend job designs.
- Snowflake Tasks and Dynamic Tables cover ELT-style orchestration cleanly but are deliberately simple; many teams pair Snowflake with an external scheduler or dbt for anything beyond linear chains.
For migrated estates with hundreds of interdependent jobs, orchestration parity — "will my 200-job dependency chain survive?" — deserves a proof of concept on whichever platform you shortlist.
Skills fit
- Databricks rewards Python/Spark engineers.
- Fabric rewards Microsoft estates: T-SQL, Visual Studio heritage, Power BI teams, existing Azure AD and Purview governance.
- Snowflake rewards SQL-literate teams and dbt practitioners, and is the gentlest landing for pure warehouse teams.
Hire-ability matters here. If your market cannot recruit Spark engineers at sensible rates — a real constraint for some UK regions — a Spark-first target has a hidden long-term cost regardless of platform merits.
4. When each platform wins
The table below is the decision matrix we use when scoping an estate; it is deliberately blunt.
| Situation | Databricks | Microsoft Fabric | Snowflake |
|---|---|---|---|
| Transformation-heavy estate (complex SSIS/Talend data flows) | Best fit | Good (Spark path) | Fair (needs SQL/dbt redesign) |
| Estate is mostly SQL-flavoured logic | Good | Best fit | Best fit |
| Existing Microsoft estate (SQL Server, Power BI, Azure AD) | Good | Best fit | Fair |
| Existing open-source / Spark skills | Best fit | Good | Fair |
| Warehouse-first analytics with BI focus | Fair | Good | Best fit |
| Complex job dependency chains | Good | Best fit | Fair (external orchestrator) |
| Cost predictability is the priority | Fair (consumption) | Good (fixed capacity) | Good (warehouse sizing) |
| Elastic, spiky batch workloads | Best fit | Fair (capacity headroom needed) | Good (auto-scaling) |
Two honest caveats. First, "best fit" in this table is about migrated ETL, not about which platform is best in the abstract — each of the three is genuinely excellent at something. Second, mixed estates exist: it is entirely reasonable to land transformation-heavy jobs in Databricks and BI-facing marts in Snowflake or Fabric, provided you accept the integration and skills overhead of running two platforms.
5. The honest way to compare costs
"Databricks vs Snowflake cost" is a search that mostly returns rate-card arithmetic, and rate-card arithmetic is close to useless for a migrated estate. Three reasons.
- The billing units are incommensurable. DBUs, capacity units and credits measure different things at different layers; dividing one by another produces a number with no meaning.
- Migrated estates behave unlike the reference workloads vendors price against. Legacy batch jobs often run at odd hours, move more data than they should, and contain inefficiencies nobody noticed because the old platform was a sunk cost. Those behaviours follow you to the new platform.
- The big cost surprises are structural, not unit-based: under-provisioned Fabric capacity causing queueing at month-end; Snowflake warehouses left running over the weekend; Databricks clusters sized by habit rather than measurement.
The defensible approach is to model cost per workload: take representative pipelines from your estate — the biggest, the nastiest, the most business-critical — rebuild or estimate them on each shortlisted platform, and observe actual consumption for a few weeks. It is more effort than reading a pricing page. It is also the only method that has ever survived contact with a finance director.
If you want a head start on the estate-level economics, the ETL migration ROI calculator models migration effort and running-cost ranges side by side.
6. The LogicLift approach
This is the one vendor section in an otherwise vendor-neutral post; judge it accordingly.
The platform question and the migration question are often answered in the wrong order. Teams commit to a platform, then discover the translation to that platform's paradigm is harder than the business case assumed. Our view is that the estate itself should inform the target: a transformation-heavy Talend estate and a SQL-flavoured SSIS estate are not the same migration, and should not land on the same platform by default.
LogicLift reads legacy packages — SSIS, Talend, Informatica, DataStage — and extracts the actual business logic rather than the orchestration shell. It then generates target-native code for whichever platform genuinely fits: PySpark notebooks for Databricks, notebooks and pipelines for Microsoft Fabric, or dbt models and SQL for a warehouse target. The same estate can be re-pointed if your platform decision changes during the programme, which removes the risk of choosing wrong in month one.
Two guardrails come as standard, whichever the target:
- Parity tests. The same synthetic inputs run through the legacy package and the generated code; outputs are diffed row by row. Equivalence is evidenced, not eyeballed.
- Human in the loop. Generated code is reviewed and signed off by data engineers before cutover. The tooling removes hand-transcription, not judgment.
For a typical estate of 100–300 packages, this lands production-ready, parity-tested code in 6–10 weeks, against 12–18 months for a manual rewrite. If you are weighing Databricks against Fabric specifically, the SSIS to Databricks page and the SSIS to Fabric migration guide show what the translated output looks like on each, and Talend estates follow the same pattern.
Choosing between Databricks, Fabric and Snowflake?
Book a free 30-minute migration assessment. We will look at your package inventory, work through the decision matrix above for your estate and team skills, and map the fastest safe route — whichever platform it lands on.
7. Real considerations before you commit
Whichever platform you choose, and whichever migration route you take, four things will be true.
- Rollback needs planning before cutover. Run legacy and new pipelines in parallel through a soak period, and require parity tests to pass on production-shaped data, not just synthetic samples. A rollback plan written after go-live is a post-mortem in waiting.
- Security and connectivity are the same problem everywhere. On-prem sources — file shares, old SQL Server instances, mainframe extracts — need self-hosted connectivity or a staged extraction regardless of target. Cloud data platforms do not magic away twenty-year-old network topology.
- Human review capacity is a real constraint. Whether you migrate manually or with assistance, senior engineers must own sign-off. Plan their time into the schedule explicitly, or the bottleneck just moves.
- Platform choice is not risk-free even when the migration goes well. Fabric capacity needs right-sizing over the first quarter; Snowflake warehouse policies need enforcing; Databricks cluster configurations need a cost review at month three. Budget a small ongoing optimisation effort, not just the migration.
8. ROI summary
- Platform choice is a ten-year decision made in a migration that lasts months. Optimise for where your team's skills and your estate's logic shape naturally fit, not for a pricing page.
- None of the three platforms wins on sticker price. Cost modelling per workload is the only honest method; list-price comparisons actively mislead.
- Manual rewrite of a 100–300 package estate takes 12–18 months and typically lands north of £1m all-in; assisted migration with parity testing lands the same estate in 6–10 weeks.
- The ROI that matters is risk-adjusted: cost of the estate to run plus the people maintaining it, versus a one-off migration with evidenced cutover. Speed to cutover belongs in that model, because a migration stalled at 80% delivers none of the benefit.
9. Where to go next
- See how LogicLift migrates SSIS to Databricks
- Read the SSIS to Fabric migration guide
- Read the Talend to Fabric migration scenario
- Model your estate with the ETL migration ROI calculator
- Book a migration assessment
- LogicLift home page
Gibran Kazi, founder of LogicLift — AI-assisted legacy ETL migration with built-in parity tests.