Foundation Model Serving
Serve state-of-the-art foundation models for both real-time and batch inference workload needs. This enables you to quickly and easily build applications that leverage high-quality generative AI models without the need to maintain your own model deployment.
* Displayed pricing does not guarantee product availability in that region. For product availability see here: AWS, Azure, GCP, SAP
1. Azure Databricks, as a first-party service on Microsoft Azure, offers unified billing and support by Microsoft
The Premium tier on Azure Databricks corresponds to the Enterprise tier on AWS and GCP
2. Hourly pricing is charged on a per-minute increment
3. Provisioned Throughput is sold in increments of model units. The amount of throughput per model unit varies by model and request shape (input tokens, output tokens, cache hit rate). Please use the GenAI Calculator to estimate workload-specific throughput and total cost
How to use DBU rates
Your final price equals the SKU price ($ per DBU, in the tiles above) multiplied by the DBU rate (DBU per million tokens or hours, from the table below).
For example, if your SKU price is $0.07 / DBU (e.g. AWS, US East N. Virginia):
- Your final input token price for Kimi K3 is:
[$0.07 / DBU] × [42.857 DBU / M input tokens] = $3.00 / M input tokens. - Your one-month Provisioned Throughput price for GLM-5.3 is:
[$0.07 / DBU] × [142.857 DBU / hour] = $10.00 / hour.
Foundation Model Serving DBU rates
Pay Per Token
| Model | Pay Per Token (DBU Per 1M Tokens) | ||
|---|---|---|---|
| Input | Output | Cache read | |
| Kimi K3 ⌖ | 42.857 | 214.286 | 4.286 |
| GLM-5.2, 5.3 | 20.000 | 62.857 | 3.714 |
| DeepSeek V4 Pro | 18.857 | 56.571 | 1.886 |
| Inkling | 14.286 | 57.857 | 2.429 |
| Kimi K2.7 | 13.571 | 57.143 | 2.714 |
| GLM-5.3 Flash | 2.143 | 7.143 | 0.429 |
| DeepSeek V4.1 Flash | 4.286 | 17.143 | 0.429 |
| DeepSeek V4 Flash | 2.000 | 4.000 | 0.400 |
| Qwen 3.5 122B | 3.143 | 31.429 | - |
| Llama 4 Maverick | 7.143 | 21.429 | - |
| Llama 3.3 70B | 7.143 | 21.429 | - |
| Qwen 3 80B Instruct | 2.143 | 17.143 | - |
| GPT-OSS-120B | 2.143 | 8.571 | - |
| Gemma 3 12B | 2.143 | 7.143 | - |
| Llama 3.1 8B | 2.143 | 6.429 | - |
| GPT-OSS-20B | 1.000 | 4.286 | - |
| GTE | 1.857 | - | - |
| BGE Large | 1.429 | - | - |
| Qwen 3 0.6B Embedding | 0.286 | - | - |
Priority Pay Per Token
| Model | Priority Pay Per Token (DBU Per 1M Tokens) | ||
|---|---|---|---|
| Input | Output | Cache read | |
| Kimi K3 ⌖ | 75.000 | 375.000 | 7.500 |
| GLM-5.2, 5.3 | 35.000 | 110.000 | 6.500 |
| GLM-5.3 Flash | 3.750 | 12.500 | 0.750 |
| Qwen 3.5 122B | 5.500 | 55.000 | - |
Provisioned Throughput
| Model | Provisioned Throughput (DBU Per Hour Per 50 Model Units) | ||
|---|---|---|---|
| On-demand | 1 month reservation | 3 month reservation | |
| Kimi K3 ⌖ | - | 142.857 | 128.571 |
| GLM-5.2, 5.3 | - | 142.857 | 128.571 |
| GLM-5.3 Flash | - | 142.857 | 128.571 |
| DeepSeek V4.1 Flash | - | 142.857 | 128.571 |
| Qwen 3.5 122B | 85.714 | - | - |
| Llama 4 Maverick | 85.714 | - | - |
| Llama 3.3 70B | 85.714 | - | - |
| Qwen 3 80B Instruct | 104.762 | - | - |
| GPT-OSS-120B | 71.429 | - | - |
| Gemma 3 12B | 71.429 | - | - |
| Llama 3.1 8B | 53.571 | - | - |
| GPT-OSS-20B | 53.571 | - | - |
| GTE | 100.000 | - | - |
| BGE Large | 24.000 | - | - |
| Qwen 3 0.6B Embedding | 50.000 | - | - |
| Llama 3.2 3B | 46.429 | - | - |
| Llama 3.2 1B | 42.857 | - | - |
Batch Inference
| Model | Batch Inference (DBU Per Hour) |
|---|---|
| Qwen 3.5 122B | 85.714 |
| Llama 4 Maverick | 85.714 |
| Llama 3.3 70B | 85.714 |
| Qwen 3 80B Instruct | 104.762 |
| GPT-OSS-120B | 71.429 |
| Gemma 3 12B | 71.429 |
| Llama 3.1 8B | 53.571 |
| GPT-OSS-20B | 53.571 |
| GTE | 100.000 |
| BGE Large | 24.000 |
| Qwen 3 0.6B Embedding | 50.000 |
| Llama 3.2 3B | 46.429 |
| Llama 3.2 1B | 42.857 |
⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).
Pay as you go with a 14-day free trial or contact us for committed-use discounts or custom requirements.