Skip to main content

Proprietary Foundation Model Serving

Serve state-of-the-art proprietary foundation models for real-time and batch inference workload needs. This enables you to quickly and easily build applications that leverage high-quality proprietary generative AI models from various vendors directly on the Databricks platform without the need to additionally and separately engage with other vendors.

Loading...

* For Azure customers, if you have an Azure Commit with Databricks, Databricks may make available this service as an ADI Service that integrates with Azure Databricks. The ADI Service is sold and invoiced by Databricks. Contact Sales to get access.
1. Azure Databricks, as a first-party service on Microsoft Azure, offers unified billing and support by Microsoft
   The Premium tier on Azure Databricks corresponds to the Enterprise tier on AWS and GCP

Proprietary Foundation Model Serving DBU rates

Pay Per Token

Model

 

Standard Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Cache write

GPT-5.6 Sol* ⌖Short context

57.143

285.714

5.714

71.429

Long context

114.286

428.571

11.429

142.857

GPT-5.6 Terra ⌖Short context

28.571

171.429

2.857

35.714

Long context

57.143

257.143

5.714

71.429

GPT-5.6 Luna ⌖Short context

2.857

17.143

0.286

3.571

Long context

5.714

25.714

0.571

7.143

GPT-5.4 Pro, 5.5 Pro ⌖Short context

428.571

2,571.43

-

-

Long context

857.143

3,857.14

-

-

GPT-5.5 ⌖Short context

71.429

428.571

7.143

-

Long context

142.857

642.857

14.286

-

GPT-5.4 ⌖Short context

35.714

214.286

3.571

-

Long context

71.429

321.429

7.143

-

GPT-5.4 Mini ⌖ 

10.714

64.286

1.071

-

GPT-5.4 Nano ⌖ 

2.857

17.857

0.286

-

GPT-5.2 Codex, 5.3 Codex ⌖ 

25

200

2.5

-

GPT-5.2 ⌖ 

25

200

2.5

-

GPT-5, 5.1 ⌖ 

17.857

142.857

1.786

-

GPT-5.1 Codex Max ⌖ 

17.857

142.857

1.786

-

GPT-5.1 Codex Mini ⌖ 

3.571

28.571

0.357

-

GPT-5 Mini ⌖ 

3.571

28.571

0.357

-

GPT-5 Nano ⌖ 

0.714

5.714

0.071

-

GPT Image 2Text tokens

71.429

-

17.857

-

Image tokens

114.286

428.571

28.571

-

GPT Image 1.5Text tokens

71.429

142.857

17.857

-

Image tokens

114.286

457.143

28.571

-

GPT Image 1Text tokens

71.429

-

17.857

-

Image tokens

142.857

571.429

35.714

-

GPT Image 1 MiniText tokens

28.571

-

2.857

-

Image tokens

35.714

114.286

3.571

-

Priority Pay Per Token

Model

 

Priority Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Cache write

GPT-5.6 Sol* ⌖Short context

114.286

571.429

11.429

142.857

Long context

228.571

857.143

22.857

285.714

GPT-5.6 Terra ⌖Short context

57.143

342.857

5.714

71.429

Long context

114.286

514.286

11.429

142.857

GPT-5.6 Luna ⌖Short context

5.714

34.286

0.571

7.143

Long context

11.429

51.429

1.143

14.286

GPT-5.5 ⌖ 

178.571

1,071.43

17.857

-

GPT-5.4 ⌖ 

71.429

428.571

7.143

-

GPT-5.4 Mini ⌖ 

21.429

128.571

2.143

-

GPT-5.3 Codex ⌖ 

50

400

5

-

GPT-5.2 ⌖ 

50

400

5

-

GPT-5, 5.1 ⌖ 

35.714

285.714

3.571

-

GPT-5 Mini ⌖ 

6.429

51.429

0.643

-

Batch Inference

Model

Batch Inference (DBU Per Hour)

GPT-5.6 Sol* ⌖

171.429

GPT-5.6 Terra ⌖

153.571

GPT-5.6 Luna ⌖

22.857

GPT-5.4 Pro, 5.5 Pro ⌖

1,142.86

GPT-5.5 ⌖

214.286

GPT-5.4 ⌖

192.857

GPT-5.4 Mini ⌖

107.143

GPT-5.4 Nano ⌖

71.429

GPT-5.2 ⌖

184.286

GPT-5, 5.1 ⌖

131.429

GPT-5 Mini ⌖

71.429

GPT-5 Nano ⌖

53.571

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

Token modalities (text, image, audio) are shown explicitly only when their numeric prices differ. Once modalities are shown, separate rows also preserve differences in supported token types; a dash means that token type is not supported for that modality. Where modalities are not shown, each displayed price applies to every modality that supports the corresponding token type.

* NOTE: The GPT-5.6 Sol rates shown here reflect OpenAI’s promotional pricing, in effect through November 21, 2026. Afterward, the standard GPT-5.6 Sol rates will take effect; input, cache, and Batch rates will be 25% higher than those shown, while output rates will be 50% higher.

Proprietary Foundation Model Serving DBU rates

Pay Per Token

Model

Standard Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Cache write

Cache write (1hr)

Claude Fable 5.1 ⌖142.857714.2863.571178.571285.714
Claude Fable 5 ⌖142.857714.28614.286178.571285.714
Claude Opus 4.5, 4.6, 4.7, 4.8, 5 ⌖71.429357.1437.14389.286142.857
Claude Opus 4, 4.1214.2861,071.4321.429267.857428.571
Claude Sonnet 5 ⌖28.571142.8572.85735.71457.143
Claude Sonnet 4.5, 4.6 ⌖42.857214.2864.28653.57185.714
Claude Sonnet 442.857214.2864.28653.571-
Claude Haiku 4.5 ⌖14.28671.4291.42917.85728.571

Batch Inference

Model

Batch Inference (DBU Per Hour)

Claude Fable 5, 5.1 ⌖

357.143

Claude Opus 4.5, 4.6, 4.7, 4.8, 5 ⌖

178.571

Claude Opus 4, 4.1

514.286

Claude Sonnet 5 ⌖

142.857

Claude Sonnet 4.5, 4.6 ⌖

214.286

Claude Sonnet 4

214.286

Claude Haiku 4.5 ⌖

114.286

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

Proprietary Foundation Model Serving DBU rates

Pay Per Token

Model

 

Standard Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Gemini 3.0 Pro, 3.1 Pro*Short context

35.714

214.286

3.571

Long context

71.429

321.429

7.143

Gemini 2.5 Pro*Short context

22.321

178.571

2.232

Long context

44.643

267.857

4.464

Gemini 3.7 Flash, 3.8 Flash** ⌖ 

10.714

53.571

1.071

Gemini 3.6 Flash** 

10.714

53.571

1.071

Gemini 3.5 Flash* ⌖ 

26.786

160.714

2.679

Gemini 3.0 Flash*Text tokens

8.929

53.571

0.893

Image tokens

8.929

-

0.893

Audio tokens

17.857

-

1.786

Gemini 2.5 Flash*Text tokens

5.357

44.643

0.536

Image tokens

5.357

-

0.536

Audio tokens

17.857

-

1.786

Gemini 3.5 Flash Lite* ⌖ 

5.357

44.643

0.536

Gemini 3.1 Flash Lite* ⌖Text tokens

4.464

26.786

0.446

Image tokens

4.464

-

0.446

Audio tokens

8.929

-

0.893

Gemini 2.5 Flash Lite*Text tokens

1.786

7.143

0.179

Image tokens

1.786

-

0.179

Audio tokens

5.357

-

0.536

Gemini 3 Pro Image* ⌖Text tokens

35.714

214.286

-

Image tokens

35.714

2,142.86

-

Gemini 3.1 Flash Image* ⌖Text tokens

8.929

53.571

-

Image tokens

8.929

1,071.43

-

Gemini 3.1 Flash Lite Image*Text tokens

4.464

26.786

-

Image tokens

4.464

535.714

-

Priority Pay Per Token

Model

 

Priority Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Gemini 3.1 Pro*Short context

64.286

385.714

6.429

Long context

128.571

578.571

12.857

Gemini 3.7 Flash, 3.8 Flash** ⌖ 

19.286

96.429

1.929

Gemini 3.6 Flash** 

19.286

96.429

1.929

Gemini 3.5 Flash* ⌖ 

48.214

289.286

4.821

Gemini 3.0 Flash*Text tokens

16.071

96.429

1.607

Image tokens

16.071

-

1.607

Audio tokens

32.143

-

3.214

Gemini 3.5 Flash Lite* ⌖ 

9.643

80.357

0.964

Gemini 3.1 Flash Lite* ⌖Text tokens

8.036

48.214

0.804

Image tokens

8.036

-

0.804

Audio tokens

16.071

-

1.607

Batch Inference

Model

Batch Inference (DBU Per Hour)

Gemini 3.0 Pro, 3.1 Pro*

230.429

Gemini 2.5 Pro*

164.286

Gemini 3.7 Flash, 3.8 Flash** ⌖

74.286

Gemini 3.6 Flash**

74.286

Gemini 3.5 Flash* ⌖

196.429

Gemini 3.0 Flash*

125

Gemini 2.5 Flash*

107.143

Gemini 3.5 Flash Lite* ⌖

121.429

Gemini 3.1 Flash Lite* ⌖

89.286

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

Token modalities (text, image, audio) are shown explicitly only when their numeric prices differ. Once modalities are shown, separate rows also preserve differences in supported token types; a dash means that token type is not supported for that modality. Where modalities are not shown, each displayed price applies to every modality that supports the corresponding token type.

* The DBU rates for these Gemini models do not include a promotional discount of 20% (promotional pricing is 20% lower than shown). The promotion expires on Jan 31, 2027, after which date all prices will revert to the DBU rates shown in this table.

** The Gemini 3.6 Flash, 3.7 Flash, and 3.8 Flash DBU rates shown here reflect a 50% promotion in effect through December 31, 2026. Afterward, the standard rates, twice those shown, will take effect.

Proprietary Foundation Model Serving DBU rates

Pay Per Token

Model

Standard Pay Per Token (DBU Per 1M Tokens)

Input

Output

Cache read

Grok 4.6* ⌖

35.714

107.143

8.929

Batch Inference

Model

Batch Inference (DBU Per Hour)

Grok 4.6* ⌖

153.571

⌖ When regional processing (data residency) is enabled, models with the ⌖ label will have a 10% uplift applied (will be charged DBU rates 10% greater than those shown in the table above).

* The DBU rates for these Grok models do not include a promotional discount of 20% (promotional pricing is 20% lower than shown). The promotion expires on Jan 31, 2027, after which date all prices will revert to the DBU rates shown in this table.

Pay as you go with a 14-day free trial or contact us for committed-use discounts or custom requirements.

Proprietary Foundation Model Serving FAQ