What Databricks has actually opened
Databricks has done substantive work to position itself as the open-lakehouse leader. Unity Catalog has been open-sourced. The Open APIs let external engines read managed tables. The Iceberg interoperability is now full and the platform supports Delta, Iceberg, Hudi, and Parquet under unified governance. The platform's own messaging is that the customer keeps one copy of data, uses any compute engine, and governs everything from one place. The engineering work is real and the architectural direction is genuinely more open than the previous-cycle platform was.
The Middle East data-platform leaders we work with are reading the architecture with a more nuanced view than the platform's marketing communicates. The open-source release of Unity Catalog and the Open APIs do reduce the storage-layer lock-in materially. The structural lock-in has moved one layer up. It now sits at the place the customer's data team interacts with the platform every day — the workflow surfaces, the governance configuration, the operational rhythm the team operates against.
This piece is the working read on where the actual lock-in sits at a typical Databricks deployment, and what the alternative architecture looks like at the multi-year horizon.
Where the lock-in actually sits
Four structural surfaces hold the customer in place even after the data-format lock-in has been removed.
The governance configuration surface. Unity Catalog's governance configuration — the access-control specification, the attribute-based access policies, the row filters, the column masks, the data-classification taxonomy — is built against the Unity Catalog data model. The customer's data team operates against this configuration daily. Migrating to a different catalogue requires translating the configuration into the alternative catalogue's data model. The institutional configuration that the team has refined over years is the asset the team would have to rebuild.
The workflow surface. The Databricks platform — the notebooks, the workflows, the SQL surface, the dashboarding tool, the broader Data Intelligence Platform user experience — is the working surface the data team operates against. The institutional muscle memory, the workflow patterns, the team's training, and the operational rhythms run against the Databricks experience. Migrating off the platform means retraining the team's working day.
The proprietary-extension dependency. Unity Catalog has Open APIs. The platform has proprietary extensions on top of the open-source core. The advanced features — the AI/BI dashboard surfaces, the Mosaic AI surfaces, the operational-intelligence surfaces — depend on the platform's proprietary capability. The customer who has built workflow against the proprietary extensions has the same lock-in dynamic at the workflow layer that the previous-cycle platform had at the data layer.
The compute-and-pricing model. The Databricks Runtime, the photon engine optimisation, the cluster-management substrate, and the broader compute-runtime advantages are real engineering improvements. The customer's workload runs faster on Databricks against the same data than against the equivalent open-stack compute. The customer pays the platform's commercial rate against that performance advantage. The migration to the open stack means rebuilding the workload against compute that performs at a different efficiency curve, and the cost-trajectory comparison needs to reflect this honestly.
What "open" actually means at the architecture level
The open storage layer is real. The Iceberg interoperability, the Delta openness, the external-engine read-and-write — each of these does what Databricks says it does. The data layer is no longer the lock-in point.
The open compute layer is partially real. The customer can run the workload on the platform of choice against the Iceberg or Delta substrate. The performance against the same data varies meaningfully across engines, and the customer's workload-design pattern that has optimised against Databricks Runtime may not perform equivalently on the alternative.
The governance configuration is increasingly open at the protocol level but the operational reality remains coupled to the platform. The Open APIs let the alternative catalogue read the policies. The data team's working day, the training, the configuration rhythm — these remain coupled to the platform.
The workflow surface remains proprietary. The notebook experience, the workflow tooling, the user-experience layer the team operates against — none of this has been open-sourced.
What the alternative architecture looks like
The alternative is the architecture we cover in adjacent pieces. The data on open formats — Iceberg, Delta — on object storage the customer operates. The catalogue on an open REST surface — the open-source Unity Catalog the enterprise self-hosts, Polaris, Apache Atlas, or the catalogue the enterprise chooses. The compute on fit-for-purpose engines — Trino, ClickHouse, DuckDB, Spark on Kubernetes — that the enterprise's platform team runs. The workflow surface on the open-source notebook and workflow tooling that the data community has converged on.
The architecture is operationally heavier in raw human cost than the managed Databricks experience. The architecture is dramatically more open at every layer. The customer's data team owns each substrate the platform investment is built on.
The customer who chooses this architecture accepts the operational overhead in exchange for the optionality, the per-workload cost control, and the institutional capability that comes with operating the open stack.
What stays with Databricks
The targeted workload where Databricks is genuinely the right tool. The ML workload where Mosaic AI delivers value the enterprise's open MLOps stack does not. The interactive notebook work where the team's operational rhythm is hard to replicate. The Spark-heavy batch workload where the Photon optimisation produces measurable cost reduction against the alternative.
The argument is not to retire Databricks. The argument is to right-size the deployment against the value the platform delivers, with the alternative architecture handling the workloads where the open stack is competitive.
The Middle East dimension
Three dimensions matter at a Middle East data-platform enterprise running Databricks.
The data-sovereignty posture. The open-architecture substrate gives the enterprise freedom to choose the hosting per workload. The Middle East sovereign-cloud direction is the architectural primary.
The institutional-capability investment. The Middle East data-engineering talent base has matured to operate the open stack. The argument for the managed-platform tradeoff has weakened.
The cost-trajectory honesty. The Databricks commercial relationship at the multi-year horizon needs to be modelled honestly against the alternative. The renewal cycle is the moment the conversation lands.
The saasinator perspective
The open-storage move from Databricks is real and material. The platform's lock-in has shifted one layer up. The enterprise that recognises where the actual lock-in sits is in the right position to negotiate the next renewal and to design the architecture against the multi-year horizon.
What to bring to the diagnostic
Bring the Databricks deployment scope, the Unity Catalog configuration breakdown, the workload-by-workload commercial commitment, and the data-sovereignty posture. The diagnostic is ten working days. The output is the architecture recommendation, the targeted-workload retention, and the open-stack migration plan for the workloads that should move. Plan a liberation at /diagnostic.