REFORGESnowflakeDataPlatform

Moving off Snowflake: a practical migration architecture for enterprise data teams

Chandrasekhar Kolarsaasinator AI10 min read

The exit is more open than the customer base realises

Snowflake itself has made the exit architecturally cleaner than it has been at any previous point. The platform natively supports Apache Iceberg tables. The Delta Direct surface lets the platform query Delta-formatted data in place without copying. The UniForm direction in the broader Lakehouse ecosystem produces Delta and Iceberg metadata against the same underlying files. The customer that has built warehouse workload against Snowflake is no longer trapped at the data layer in the way the previous-generation closed-format vendors required.

The structural argument for the exit changes accordingly. The argument is no longer about the technical feasibility. The argument is about the multi-year commercial trajectory, the optionality at the next renewal, and the institutional capability the data team builds in operating the alternative architecture. Data teams running on Snowflake at meaningful scale are arriving at the exit conversation with this read.

This piece is the working migration architecture for the data leader running that conversation.

What the migration architecture actually contains

The architecture has five structural components.

The storage layer. The data moves to open formats — Iceberg, Delta, the broader open-table ecosystem — on object storage the customer operates. The platform team chooses the hosting against the data-sovereignty posture the enterprise has agreed with the regulator. The data is portable across cloud providers and queryable from multiple engines.

The catalogue layer. The metadata, the table definitions, the schema-evolution history, the partition-and-clustering specification, and the access-control policies live in an open REST catalog — Iceberg REST, Unity Catalog in its open variant, Polaris, the broader open-catalogue ecosystem. The catalogue is the substrate every engine reads from.

The compute layer. The query workload runs on the engines the customer chooses for each workload class. Trino or StarRocks for ad-hoc SQL across large tables. DuckDB inside applications for the small-to-medium analytical work at the lowest latency. ClickHouse for high-cardinality OLAP. Apache Spark for the ML and large-batch work. Each engine reads from the open catalogue. The compute is fit-for-purpose per workload.

The semantic and BI layer. The metric definitions live in an open semantic layer — Cube, dbt semantic layer, Malloy. The BI tools — the enterprise's choice from the broad available catalogue — read from the semantic layer. The metric definitions are portable.

The data-engineering and orchestration layer. The pipelines run on Airflow, Dagster, or Prefect against the open formats. The data quality, the lineage tracking, the observability run on the open-source MLOps and DataOps tools the data team operates.

The customer's data team owns each layer. The vendor relationships are decoupled from the architecture choice at each layer. Switching the BI tool is not a data-migration project. Switching the compute engine is not a data-rewrite project. Switching the cloud provider is a data-replication project against the same format.

The migration sequence

The migration runs as a phased move over two to four quarters for a typical enterprise. The Snowflake workload continues during the parallel period.

Quarter one — write twice. Every new table that enters the warehouse is written to both Snowflake and to Iceberg on object storage. The platform team builds the dual-write pipeline. The Snowflake queries continue against the Snowflake-native tables. The new workloads start reading from the Iceberg surface.

Quarter two — new workloads on open. New analytical workloads, new dashboards, new ML pipelines, new ad-hoc queries — all read from the Iceberg surface against the open catalogue. The Snowflake workload continues for the existing reports and the existing pipelines. The data team builds operational confidence on the open stack.

Quarter three — heavy-query migration. The Snowflake queries that consume the most credit are identified and moved to the open compute. Trino, ClickHouse, or the engine that fits the workload class handles the workload against the same Iceberg data. The Snowflake credit consumption bends.

Quarter four — surface retirement. The Snowflake workload that remains is the workload where Snowflake is genuinely the right engine. The renewal at the end of the cycle reflects the smaller footprint. The Snowflake relationship continues at the targeted scope.

What stays on Snowflake

The targeted workload that Snowflake is competent at. The ad-hoc analytical queries with unpredictable concurrency, the cross-tenant data sharing the enterprise has built into the Snowflake Marketplace, the Snowflake-native applications the data team has invested in. The platform continues to operate against the targeted scope. The argument is not to retire the platform. The argument is to right-size the relationship.

The AI workload that has built against Cortex moves to the architecture covered separately in our Cortex piece. The portability strategy at the AI layer is independent of the warehouse-workload migration.

What changes operationally

Three things change for the enterprise running the migration.

The consumption trajectory bends. The Snowflake credit consumption reduces in proportion to the workload that has moved. The open-stack cost is the customer's own infrastructure cost, which scales differently from the consumption-based meter.

The institutional capability matures. The data team owns the substrate. The next decision — the BI tool, the compute engine, the cloud provider — is the data team's decision, not the vendor's.

The renewal trajectory changes. The Snowflake commercial conversation at the next renewal lands against a customer who has demonstrated the alternative. The targeted scope is real and the leverage is real.

The regional dimension

Three dimensions matter at a Middle East enterprise.

The data-sovereignty posture. The open-stack architecture lets the enterprise choose the hosting per workload. The UAE Sovereign Cloud direction, the KSA national-cloud strategy, the broader regional data-residency framework — each is the regulator's expectation against the enterprise's architecture. The open stack respects this through construction.

The cost trajectory at scale. The data-platform enterprises at the largest scale have seen the Snowflake consumption compound across renewal cycles in a way that the CFO modelling at signing did not anticipate. The migration architecture changes the trajectory.

The talent ecosystem has matured. The open-stack operation requires the engineering capability the regional data-engineering talent base has been building. The argument for the managed-service tradeoff against the open-stack architecture has weakened.

The saasinator perspective

Snowflake is competent at its targeted workload. The migration architecture above is not the argument that the platform should be retired. The argument is that the customer's data architecture should be open by construction, with the platform handling the workloads it is competent at and the open stack handling everything else.

The data leader who runs the first quarter of the migration successfully has changed the institutional default. The next quarter is a smaller decision.

What to bring to the diagnostic

Bring the current Snowflake deployment scope, six months of credit consumption against the workload breakdown, the BI-and-analytics inventory, and the data-sovereignty commitments. The diagnostic is ten working days. The output is the migration architecture, the first-quarter scope, and the workload-priority recommendation. Book a diagnostic at /diagnostic.


Share this insight