Foundation Model Serving
Serve state-of-the-art foundation models for both real-time and batch inference workload needs. This enables you to quickly and easily build applications that leverage high-quality generative AI models without the need to maintain your own model deployment.
* Displayed pricing does not guarantee product availability in that region. For product availability see here: AWS, Azure, GCP, SAP
1. Azure Databricks, as a first-party service on Microsoft Azure, offers unified billing and support by Microsoft
The Premium tier on Azure Databricks corresponds to the Enterprise tier on AWS and GCP
2. Hourly pricing is charged on a per-minute increment
3. Throughput in a single unit of PT capacity varies by model and query shape (input vs. output tokens). Please use the GenAI Calculator to estimate workload-specific throughput and total cost
Foundation Model Serving DBU rates
Pay Per Token
Model | Standard Pay Per Token (DBU Per 1M Tokens) | ||
|---|---|---|---|
Input | Output | Cache read | |
| Kimi K3 ⌖ | 42.857 | 214.286 | 4.286 |
| GLM-5.2, 5.3 | 20.000 | 62.857 | 3.714 |
| DeepSeek V4 Pro | 18.857 | 56.571 | 1.886 |
| Inkling | 14.286 | 57.857 | 2.429 |
| Kimi K2.7 | 13.571 | 57.143 | 2.714 |
| GLM-5.3 Flash | 2.143 | 7.143 | 0.429 |
| DeepSeek V4 Flash | 2.000 | 4.000 | 0.400 |
| Qwen 3.5 122B | 3.143 | 31.429 | - |
| Llama 4 Maverick | 7.143 | 21.429 | - |
| Llama 3.3 70B | 7.143 | 21.429 | - |
| Qwen 3 80B Instruct | 2.143 | 17.143 | - |
| GPT-OSS-120B | 2.143 | 8.571 | - |
| Gemma 3 12B | 2.143 | 7.143 | - |
| Llama 3.1 8B | 2.143 | 6.429 | - |
| GPT-OSS-20B | 1.000 | 4.286 | - |
| GTE | 1.857 | - | - |
| BGE Large | 1.429 | - | - |
| Qwen 3 0.6B Embedding | 0.286 | - | - |
Priority Pay Per Token
Model | Priority Pay Per Token (DBU Per 1M Tokens) | ||
|---|---|---|---|
Input | Output | Cache read | |
| GLM-5.2 | 35.000 | 110.000 | 6.500 |
| Qwen 3.5 122B | 6.286 | 62.857 | - |
Provisioned Throughput
Model | Provisioned Throughput (DBU Per Hour) | ||
|---|---|---|---|
On-demand | 1 month reservation | 3 month reservation | |
| GLM-5.2 | - | 142.857 | 121.429 |
| Qwen 3.5 122B | 85.714 | - | - |
| Llama 4 Maverick | 85.714 | - | - |
| Llama 3.3 70B | 85.714 | - | - |
| Qwen 3 80B Instruct | 78.571 | - | - |
| GPT-OSS-120B | 71.429 | - | - |
| Gemma 3 12B | 71.429 | - | - |
| Llama 3.1 8B | 53.571 | - | - |
| GPT-OSS-20B | 53.571 | - | - |
| GTE | 20.000 | - | - |
| BGE Large | 24.000 | - | - |
| Qwen 3 0.6B Embedding | 25.000 | - | - |
| Llama 3.2 3B | 46.429 | - | - |
| Llama 3.2 1B | 42.857 | - | - |
Batch Inference
Model | Batch Inference (DBU Per Hour) |
|---|---|
| Qwen 3.5 122B | 85.714 |
| Llama 4 Maverick | 85.714 |
| Llama 3.3 70B | 85.714 |
| Qwen 3 80B Instruct | 78.571 |
| GPT-OSS-120B | 71.429 |
| Gemma 3 12B | 71.429 |
| Llama 3.1 8B | 53.571 |
| GPT-OSS-20B | 53.571 |
| GTE | 20.000 |
| BGE Large | 24.000 |
| Qwen 3 0.6B Embedding | 25.000 |
| Llama 3.2 3B | 46.429 |
| Llama 3.2 1B | 42.857 |
⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).
Pay as you go with a 14-day free trial or contact us for committed-use discounts or custom requirements.