Skip to main content

Foundation Model Serving

Serve state-of-the-art foundation models for both real-time and batch inference workload needs. This enables you to quickly and easily build applications that leverage high-quality generative AI models without the need to maintain your own model deployment.

Loading...

* Displayed pricing does not guarantee product availability in that region. For product availability see here: AWS, Azure, GCP, SAP
1. Azure Databricks, as a first-party service on Microsoft Azure, offers unified billing and support by Microsoft
   The Premium tier on Azure Databricks corresponds to the Enterprise tier on AWS and GCP
2. Hourly pricing is charged on a per-minute increment
3. Provisioned Throughput is sold in increments of model units. The amount of throughput per model unit varies by model and request shape (input tokens, output tokens, cache hit rate). Please use the GenAI Calculator to estimate workload-specific throughput and total cost

How to use DBU rates

Your final price equals the SKU price ($ per DBU, in the tiles above) multiplied by the DBU rate (DBU per million tokens or hours, from the table below).

For example, if your SKU price is $0.07 / DBU (e.g. AWS, US East N. Virginia):

  • Your final input token price for Kimi K3 is:
    [$0.07 / DBU] × [42.857 DBU / M input tokens] = $3.00 / M input tokens.
  • Your one-month Provisioned Throughput price for GLM-5.3 is:
    [$0.07 / DBU] × [142.857 DBU / hour] = $10.00 / hour.

Foundation Model Serving DBU rates

Pay Per Token

ModelPay Per Token (DBU Per 1M Tokens)
InputOutputCache read
Kimi K3 ⌖42.857214.2864.286
GLM-5.2, 5.320.00062.8573.714
DeepSeek V4 Pro18.85756.5711.886
Inkling14.28657.8572.429
Kimi K2.713.57157.1432.714
GLM-5.3 Flash2.1437.1430.429
DeepSeek V4.1 Flash4.28617.1430.429
DeepSeek V4 Flash2.0004.0000.400
Qwen 3.5 122B3.14331.429-
Llama 4 Maverick7.14321.429-
Llama 3.3 70B7.14321.429-
Qwen 3 80B Instruct2.14317.143-
GPT-OSS-120B2.1438.571-
Gemma 3 12B2.1437.143-
Llama 3.1 8B2.1436.429-
GPT-OSS-20B1.0004.286-
GTE1.857--
BGE Large1.429--
Qwen 3 0.6B Embedding0.286--

Priority Pay Per Token

ModelPriority Pay Per Token (DBU Per 1M Tokens)
InputOutputCache read
Kimi K3 ⌖75.000375.0007.500
GLM-5.2, 5.335.000110.0006.500
GLM-5.3 Flash3.75012.5000.750
Qwen 3.5 122B5.50055.000-

Provisioned Throughput

ModelProvisioned Throughput (DBU Per Hour Per 50 Model Units)
On-demand1 month reservation3 month reservation
Kimi K3 ⌖-142.857128.571
GLM-5.2, 5.3-142.857128.571
GLM-5.3 Flash-142.857128.571
DeepSeek V4.1 Flash-142.857128.571
Qwen 3.5 122B85.714--
Llama 4 Maverick85.714--
Llama 3.3 70B85.714--
Qwen 3 80B Instruct104.762--
GPT-OSS-120B71.429--
Gemma 3 12B71.429--
Llama 3.1 8B53.571--
GPT-OSS-20B53.571--
GTE100.000--
BGE Large24.000--
Qwen 3 0.6B Embedding50.000--
Llama 3.2 3B46.429--
Llama 3.2 1B42.857--

Batch Inference

ModelBatch Inference (DBU Per Hour)
Qwen 3.5 122B85.714
Llama 4 Maverick85.714
Llama 3.3 70B85.714
Qwen 3 80B Instruct104.762
GPT-OSS-120B71.429
Gemma 3 12B71.429
Llama 3.1 8B53.571
GPT-OSS-20B53.571
GTE100.000
BGE Large24.000
Qwen 3 0.6B Embedding50.000
Llama 3.2 3B46.429
Llama 3.2 1B42.857

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

Pay as you go with a 14-day free trial or contact us for committed-use discounts or custom requirements.

Foundation Model Serving FAQ