Skip to main content

Foundation Model Serving

Serve state-of-the-art foundation models for both real-time and batch inference workload needs. This enables you to quickly and easily build applications that leverage high-quality generative AI models without the need to maintain your own model deployment.

Loading...

* Displayed pricing does not guarantee product availability in that region. For product availability see here: AWSAzureGCPSAP
1. Azure Databricks, as a first-party service on Microsoft Azure, offers unified billing and support by Microsoft
   The Premium tier on Azure Databricks corresponds to the Enterprise tier on AWS and GCP
2. Hourly pricing is charged on a per-minute increment
3. Throughput in a single unit of PT capacity varies by model and query shape (input vs. output tokens). Please use the GenAI Calculator to estimate workload-specific throughput and total cost

Foundation Model Serving DBU rates

Pay Per Token

Model

Standard Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Kimi K3 ⌖

42.857

214.286

4.286

GLM-5.2, 5.3

20.000

62.857

3.714

DeepSeek V4 Pro

18.857

56.571

1.886

Inkling

14.286

57.857

2.429

Kimi K2.7

13.571

57.143

2.714

GLM-5.3 Flash

2.143

7.143

0.429

DeepSeek V4 Flash

2.000

4.000

0.400

Qwen 3.5 122B

3.143

31.429

-

Llama 4 Maverick

7.143

21.429

-

Llama 3.3 70B

7.143

21.429

-

Qwen 3 80B Instruct

2.143

17.143

-

GPT-OSS-120B

2.143

8.571

-

Gemma 3 12B

2.143

7.143

-

Llama 3.1 8B

2.143

6.429

-

GPT-OSS-20B

1.000

4.286

-

GTE

1.857

-

-

BGE Large

1.429

-

-

Qwen 3 0.6B Embedding

0.286

-

-

Priority Pay Per Token

Model

Priority Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

GLM-5.2

35.000

110.000

6.500

Qwen 3.5 122B

6.286

62.857

-

Provisioned Throughput

Model

Provisioned Throughput (DBU Per Hour)

On-demand

1 month reservation

3 month reservation

GLM-5.2

-

142.857

121.429

Qwen 3.5 122B

85.714

-

-

Llama 4 Maverick

85.714

-

-

Llama 3.3 70B

85.714

-

-

Qwen 3 80B Instruct

78.571

-

-

GPT-OSS-120B

71.429

-

-

Gemma 3 12B

71.429

-

-

Llama 3.1 8B

53.571

-

-

GPT-OSS-20B

53.571

-

-

GTE

20.000

-

-

BGE Large

24.000

-

-

Qwen 3 0.6B Embedding

25.000

-

-

Llama 3.2 3B

46.429

-

-

Llama 3.2 1B

42.857

-

-

Batch Inference

Model

Batch Inference (DBU Per Hour)

Qwen 3.5 122B

85.714

Llama 4 Maverick

85.714

Llama 3.3 70B

85.714

Qwen 3 80B Instruct

78.571

GPT-OSS-120B

71.429

Gemma 3 12B

71.429

Llama 3.1 8B

53.571

GPT-OSS-20B

53.571

GTE

20.000

BGE Large

24.000

Qwen 3 0.6B Embedding

25.000

Llama 3.2 3B

46.429

Llama 3.2 1B

42.857

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

Pay as you go with a 14-day free trial or contact us for committed-use discounts or custom requirements.

Foundation Model Serving FAQ