What Cortex actually is
Snowflake Cortex AI is the fully managed generative AI layer built into the Snowflake Data Cloud. The architecture sits inside the warehouse perimeter, with the LLM functions, the embedding functions, the vector search, and the ML functions all operating against the customer's data without moving it out of the platform. The model catalogue is broad — Mistral, Llama, Snowflake Arctic, Reka, Google Gemma — and the per-call commercial model is the standard consumption-based unit on top of the warehouse compute the customer already pays for.
The pitch is internally consistent. The architecture is the warehouse becoming the intelligence layer. The integration friction with the data is by definition zero because the AI runs against the data in place. The vector store is native to the platform rather than a separate piece of infrastructure. The model governance, the access control, and the audit trail run against the same Snowflake controls the customer's data team already operates.
The Middle East data-platform leaders that we work with have a more nuanced read on the architecture. The engineering is competent. The Cortex commercial model compounds against the warehouse commercial model. The lock-in is structural in a way the headline pricing does not surface, and the conversation worth having is what the lock-in actually costs at the three-renewal-cycle horizon.
What the lock-in actually contains
The Cortex architecture creates lock-in across four structural surfaces that the headline pricing does not directly reflect.
The ML model coupling. The customer that builds an ML workflow against Cortex's ML functions has built the model logic inside the Snowflake runtime. The model is deployed against the warehouse compute. The serving is structural to the platform. Migrating the model out of Snowflake means rebuilding the model deployment against a different runtime, against different governance, against a different serving infrastructure.
The vector-store coupling. The Cortex vector search is native to the platform. The customer that builds a retrieval-augmented workflow against the Cortex vector store has the embeddings, the index, and the retrieval logic all inside Snowflake. Migrating the vector workload out means re-embedding, re-indexing, and re-implementing the retrieval logic against a different vector platform.
The LLM-call coupling. The Cortex LLM functions abstract the model choice through Snowflake's catalogue. The application logic that calls the LLM is written against the Cortex API surface. Migrating the workload to a different LLM provider — Anthropic API, OpenAI API, the Bedrock or Vertex equivalents — requires rewriting against a different API contract. The lock-in is at the API-surface level, not just at the data level.
The commercial-model compounding. The Cortex consumption sits on top of the warehouse consumption. The customer that has built workflow against Cortex pays the Cortex compute against every call, and the warehouse compute against the data the call reads. The total cost trajectory across the three-renewal horizon compounds against both meters.
What the alternative architecture looks like
We do not propose retiring Snowflake at every customer that has invested in the platform. The warehouse is competent at the warehouse workload and the migration friction is real. The conversation worth having is which of the AI workloads the customer should own outright on infrastructure the customer operates, with the Snowflake warehouse continuing to handle the warehouse workload.
The pattern is the architecture we run in our broader data-platform pieces. The data layer migrates to open formats — Iceberg, Delta — on object storage the customer operates. The warehouse compute continues on Snowflake for the workloads where Snowflake is competent. The AI workloads — the LLM calls, the embedding generation, the vector retrieval, the ML model deployment — move to infrastructure the customer operates against the model layer the customer chooses.
The ML workflow runs on the customer's own MLOps stack. MLflow, Kubeflow, the broader open-MLOps ecosystem — all of these work against the open-format data on object storage. The model serving is the customer's, the eval suite is the customer's, the deployment cadence is the customer's.
The vector workflow runs on the customer's own vector platform. pgvector, Qdrant, Weaviate, Milvus — each of these is an open-source-or-permissively-licensed option that the customer operates. The embeddings are the customer's, the index is the customer's, the retrieval logic is the customer's.
The LLM-call workflow runs against the model layer the customer chooses. Anthropic, OpenAI, the open-weight models the customer deploys on owned infrastructure, the regional model providers — each is a structured choice the customer makes at the architecture level. The per-call cost is decoupled from the warehouse compute.
The portability strategy
The customer who runs the alternative architecture builds the portability strategy as the first principle. Three structural patterns matter.
The data-format independence. The open formats — Iceberg, Delta — are the customer's substrate. The warehouse compute reads from them. The AI workloads read from them. The cross-engine read is the architectural primary.
The model-API independence. The LLM-call workflow goes through the customer's own model-routing layer. The application logic calls the customer's API. The customer's API routes to the model provider the customer has selected for the workload. Switching the underlying provider is a routing change, not an application rewrite.
The vector-platform independence. The embeddings are stored in the customer's vector platform, with the embedding-generation workflow decoupled from the storage layer. Switching the underlying vector platform is a migration of the index against the same embeddings.
What stays with Snowflake
The warehouse workload that Snowflake is competent at stays. The ad-hoc analytical queries, the BI workload, the cross-tenant data sharing on the Snowflake Marketplace where the customer is exposed to external data products, the warehouse-native applications that the customer has built and validated. The argument is not against Snowflake. The argument is that the AI workloads have a portability requirement the warehouse workloads do not, and the architecture should reflect that.
The Middle East dimension
Three dimensions matter at a Middle East data-platform customer.
The data-sovereignty posture. The Middle East regulators have published expectations on data residency and on cloud-provider concentration. The portability architecture lets the enterprise choose the hosting per workload against the specific regulatory expectation.
The AI workload growth trajectory. The Vision 2030 and UAE strategic-direction commitments to AI-native operating models are driving the AI workload at Middle East enterprises faster than the warehouse workload is growing. The portability architecture handles the asymmetric growth.
The institutional capability. The Middle East data-engineering talent base has matured to operate the open-MLOps and the model-routing layer. The argument that the customer needs the managed-AI service to compensate for talent gaps no longer applies at most enterprises.
The saasinator perspective
Cortex is competent engineering. The lock-in is real. The architecture worth running is the one where the AI workload is portable by construction, regardless of where the warehouse compute runs.
The customer who builds the portability strategy is in a different commercial conversation with Snowflake at every subsequent renewal. The Cortex usage is targeted at the workloads where the in-warehouse value is clearest, and the rest runs on the architecture the customer operates.
What to bring to the diagnostic
Bring the current Snowflake deployment scope, the Cortex usage breakdown, the AI workload roadmap, and the data-sovereignty commitments. The diagnostic is ten working days. The output is the architecture recommendation and the first-quarter portability scope. Book a diagnostic at /diagnostic.