PLATFORM · DATABRICKS · MOSAIC AI

Mosaic AI per token
on top of DBUs per Spark job.

Foundation Model APIs, Model Serving, Pretraining — every layer adds a meter over the credits you already burn. You do not have to open by cancelling any of it. Start by building the same stack beside it.

  • You own the models and the evals
  • No per-token Mosaic premium
  • Runs where you choose
The estatereplaceable
  • Foundation Model APIsTraining pipeline
  • Mosaic AI Model ServingServing layer
  • Mosaic AI PretrainingRetrieval index
  • Databricks Vector SearchAgent runtime
  • Mosaic AI Agent Framework · GenieAnalytics surface
  • AI/BI DashboardsMLOps audit trail
Foundation Model APIs · Model Serving · Vector Search · Agent Frameworkyours on commit one
Why this domain exists

The technology underneath is open.
The packaging around it is not.

The capability is real — much of it open source you could run yourself. Around it sits a per-token meter attached to the consumption commitment already carrying the lakehouse.

Mosaic AI is the GenAI tier on the Databricks Lakehouse — Foundation Model APIs, Model Serving, Pretraining and Fine-tuning, Vector Search, Agent Framework and Genie. Every layer adds a per-token or per-endpoint meter on top of the DBU credits already running the lakehouse, while most of the technology underneath — MLflow, Ray, vLLM, the Llama family — is open source.

Priced per token, per endpoint

Foundation Model APIs, Mosaic AI Model Serving and Mosaic AI Pretraining each meter on top of the DBU credits already running the lakehouse. Training more grows the bill twice.

Stacked on consumption you already pay

Vector Search, Agent Framework, Genie and AI/BI Dashboards each add a tier over the same platform bill. Small on their own; together a second platform budget.

Open underneath, licensed on top

MLflow, Ray, vLLM and the Llama family are open source. The premium sits in the packaging and the runtime, not in technology you could not otherwise run.

Double meterDatabricks' published pricing model. Both meters are theirs, not ours.
Per DBU, then per tokenTraining more grows the bill twice.
A job run on the lakehouse you already commit to
The DBU credits already running the lakehouse
Foundation Model APIs and Mosaic AI metering on top
STAGE 01Ignite

Build the stack beside the one you meter.

Serving, retrieval and agent orchestration of your own, running against the same lakehouse, next to Mosaic AI. Nothing retired, nothing migrated — the Databricks contract is untouched on day one.

Model serving

Production inference on KServe, BentoML, vLLM or TensorRT-LLM, running on infrastructure your platform team operates and can scale on its own terms.

Runs beside Mosaic AI Model Serving

Vector retrieval

Vector retrieval on pgvector, Qdrant, Weaviate or Vespa against the lakehouse you already run, with indexes held where you choose.

Runs beside Databricks Vector Search

Agent platform

Multi-step agentic workflows on the orchestration framework you pick, reading your lakehouse and writing back when a human signs off.

Runs beside Mosaic AI Agent Framework and Genie

Model choice

Pick the foundation model that fits the workload and swap it without a vendor migration. Mosaic AI ships Databricks' model selection; you ship yours.

Runs beside Foundation Model APIs

STAGE 02Reforge

What your platform team already built on Mosaic AI, rebuilt as software you own.

The training and fine-tuning runs your engineers run, the natural-language analytics your analysts built, and the MLOps records under both. Same logic, modern substrate, and the per-token meter stops.

IN — what you run today
  • Foundation Model APIs
  • Mosaic AI Model Serving
  • Mosaic AI Pretraining
  • Databricks Vector Search
  • Mosaic AI Agent Framework · Genie
  • AI/BI Dashboards
SAIF

the saasinator AI Factory — glass-walled delivery

  • Brief
  • Build
  • Evals
  • Deploy
  • Transfer
OUT — what you own afterwards
  • Training pipeline
  • Serving layer
  • Retrieval index
  • Agent runtime
  • Analytics surface
  • MLOps audit trail

Product names are Databricks' own. What comes out the other side is yours — source, models, prompts, evals and pipeline, transferred on commit one.

Training and fine-tuning rebuild

The pretraining and fine-tuning runs your engineers already operate, rebuilt on Ray, Kubeflow and MLflow running on infrastructure you control.

Mosaic AI Pretraining · Foundation Model Fine-tuning

Analytics surface rebuild

The natural-language analytics your analysts already built, rebuilt on open tooling against your lakehouse and gated by an eval suite your team can re-run.

AI/BI Dashboards · Genie

MLOps audit rebuild

Every training run, model deploy and inference call logged in a registry your platform team operates, so evidence does not sit behind the platform meter.

MLflow on the Databricks Lakehouse

Coverage

Four surfaces. The same pattern in each.

Wherever token volume is highest is where the per-token meter bites hardest. These are the surfaces we rebuild, and what sits inside each.

Training
  • Pretraining and fine-tuning on Ray, Kubeflow and MLflow
  • Runs on infrastructure your platform team controls
  • Checkpoints and datasets held as assets you own
Serving
  • Production inference on KServe, BentoML, vLLM or TensorRT-LLM
  • Endpoints scaled on your terms, not per-endpoint pricing
  • Model choice recorded per workload
Retrieval & agents
  • Vector retrieval on pgvector, Qdrant, Weaviate or Vespa
  • Multi-step agents reading your lakehouse and calling tools
  • No per-index fee for retrieving your own data
Analytics & governance
  • Natural-language analytics on open tooling against your lakehouse
  • Every training run, deploy and inference call logged
  • Registry your platform team operates and can inspect
Functional areas this domain touches
STAGE 03Liberate

The tier comes out. The workloads stay — as software you own.

Liberation is earned, not sold. By the time Mosaic AI comes off the DBU bill, you have watched us build the stack and run the evals that gate it.

The licence ledger

Illustrative

The commercial shape of a GenAI tier, as a buyer reads it

Basis of charge
Per token and per endpoint, metered on top of the consumption credits already running the platform. Usage is the meter twice over.
Discount condition
Conditional on the consumption commitment already made for the platform, not on the AI tier standing on its own merits.
Tier structure
Serving, training, retrieval, orchestration and conversational analytics each price separately on top of the same platform bill.
On exit
Training runs, registries, indexes and agent definitions sit inside the vendor's platform — not in an asset you hold.
Cost of staying————

Shape only — the direction of travel, not a quantity. Your own curve comes from your own consumption commitment and your own token volume.

Illustrative. This is our reading of a commercial pattern, not a quotation from any Databricks agreement — no contract text, no clause references, no figures.

Foundation Model APIs replacement

Inference and serving run on open MLOps you operate, so the per-token and per-endpoint lines come off the platform bill entirely.

Foundation Model APIs · Mosaic AI Model Serving

Agent Framework replacement

Agentic orchestration owned outright, retired into only once the replacement is carrying the work.

Mosaic AI Agent Framework · Genie

Vector Search replacement

Retrieval indexes held on your own infrastructure, so searching your own lakehouse is not charged per index.

Databricks Vector Search

How it is built

Glass-walled from brief to transfer. Nothing behind a black box.

SAIF is our delivery method and it runs in the open. You watch the build as it happens, read the evals that gate every release, and keep every artefact — including the ones that record what did not work.

  1. 01

    Brief

    One workflow, scoped against your own data and your own renewal position.

  2. 02

    Build

    Agentic delivery against your systems, visible while it runs.

  3. 03

    Evals

    Every release gated on tests you can read and re-run yourself.

  4. 04

    Deploy

    Into infrastructure you control, alongside the system it stands beside.

  5. 05

    Transfer

    Your team runs it. We do not leave until they can.

You own it from commit one

Source, models, prompts, evals and pipeline. Not a licence to use what we built — the asset itself.

Two weeks to a working build

A working build against your own systems in two weeks. Fixed scope, fixed bill.

It runs beside the Mosaic AI service first

Nothing is retired on faith. The Mosaic AI service goes when the replacement is carrying the work.

Proof

What we will put in writing.

100%
IP transferred on commit one
0
per-token Mosaic credit consumption
2 weeks
to a working build · fixed scope, fixed bill
Our commitment

Every engagement starts with a scoped working build against your own systems. If it doesn't convince you, you pay nothing — and you keep the code either way.

The ask

Bring one Mosaic AI use case and your consumption commitment.

Ten working days. Which workload to build first, what the per-token meter adds to the DBU bill, and what owning the same workload costs to build. You keep the analysis.

Fixed feeTen working daysNo commitment beyond the diagnostic