Azure Speed Test

Prices by deployment type

All rates are USD per 1 million tokens. Each row uses its own deployment type and published price regions. A range reflects regional differences within that type. Missing prices mean no published meter in this snapshot; they do not mean free usage. A retail price does not guarantee deployment availability or quota.

GPT-5 mini OpenAI prices in USD per 1 million tokens.
Deployment typeInput / 1MOutput / 1MCached input / 1MCache write / 1MPrice regions
Data Zone Batch $0.1375 $1.10 $0.01375 Not published 12
Data Zone Priority $0.495 $3.96 $0.0495 Not published 14
Data Zone Standard $0.275 $2.20 $0.0275 Not published 15
Global Batch $0.125 $1.00 $0.0125 Not published 23
Global Priority $0.45 $3.60 $0.045 Not published 25
Global Standard $0.25 $2.00 $0.025 Not published 27

The tables contain only this OpenAI listing. Other providers, versions and similarly named Azure products can have different meters. Some token categories may be missing in individual regions; expand the regional tables for exact coverage.

How token costs are calculated

For each region and deployment type, multiply each token category by its matching rate, then divide the sum by 1,000,000. Monthly estimates multiply that request cost by the number of requests. Ranges use complete workloads priced in each region, so input and output prices from different regions are never combined.

Request cost = (uncached input tokens × input rate + cached input tokens × cached input rate + cache-write input tokens × cache-write rate + output tokens × output rate) / 1,000,000.

The examples below use zero cache writes. The cached example assumes half of the input tokens are eligible cache reads on every request; actual cache hit rates and any cache creation charges depend on the deployment. Missing cache rates are not replaced with ordinary input rates. Estimates cover these token meters only, before taxes, negotiated discounts and any separate hosting, tools, media or provisioned-capacity charges.

Data Zone Batch prices and cost examples

Data Zone processing
Inference data is processed within the Microsoft-specified US, EU, or APAC data zone.
Asynchronous Batch
Batch meters cover asynchronous jobs. Microsoft documents Global and Data Zone Batch at a 50% discount, with a 24-hour target for Global Batch; model support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.0015125
Per month
$15.125

Across 12 published price regions.

Customize workload or price region for Data Zone Batch, No cached input

50% cached input

Per request
$0.00132688
Per month
$13.26875

Across 12 published price regions.

Customize workload or price region for Data Zone Batch, 50% cached input
View 12 region prices for Data Zone Batch

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Data Zone Batch exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
East USeastus$0.1375 $1.10 $0.01375 Not published
East US 2eastus2$0.1375 $1.10 $0.01375 Not published
France Centralfrancecentral$0.1375 $1.10 $0.01375 Not published
Germany West Centralgermanywestcentral$0.1375 $1.10 $0.01375 Not published
North Central USnorthcentralus$0.1375 $1.10 $0.01375 Not published
Poland Centralpolandcentral$0.1375 $1.10 $0.01375 Not published
South Central USsouthcentralus$0.1375 $1.10 $0.01375 Not published
Spain Centralspaincentral$0.1375 $1.10 $0.01375 Not published
Sweden Centralswedencentral$0.1375 $1.10 $0.01375 Not published
West Europewesteurope$0.1375 $1.10 $0.01375 Not published
West USwestus$0.1375 $1.10 $0.01375 Not published
West US 3westus3$0.1375 $1.10 $0.01375 Not published

Data Zone Priority prices and cost examples

Data Zone processing
Inference data is processed within the Microsoft-specified US, EU, or APAC data zone.
Priority pay-per-token
Priority processing uses pay-as-you-go token meters for faster responses. Model and deployment support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.005445
Per month
$54.45

Across 14 published price regions.

Customize workload or price region for Data Zone Priority, No cached input

50% cached input

Per request
$0.00477675
Per month
$47.7675

Across 14 published price regions.

Customize workload or price region for Data Zone Priority, 50% cached input
View 14 region prices for Data Zone Priority

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Data Zone Priority exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Central UScentralus$0.495 $3.96 $0.0495 Not published
East USeastus$0.495 $3.96 $0.0495 Not published
East US 2eastus2$0.495 $3.96 $0.0495 Not published
France Centralfrancecentral$0.495 $3.96 $0.0495 Not published
Germany West Centralgermanywestcentral$0.495 $3.96 $0.0495 Not published
North Central USnorthcentralus$0.495 $3.96 $0.0495 Not published
North Europenortheurope$0.495 $3.96 $0.0495 Not published
Poland Centralpolandcentral$0.495 $3.96 $0.0495 Not published
South Central USsouthcentralus$0.495 $3.96 $0.0495 Not published
Spain Centralspaincentral$0.495 $3.96 $0.0495 Not published
Sweden Centralswedencentral$0.495 $3.96 $0.0495 Not published
West Europewesteurope$0.495 $3.96 $0.0495 Not published
West USwestus$0.495 $3.96 $0.0495 Not published
West US 3westus3$0.495 $3.96 $0.0495 Not published

Data Zone Standard prices and cost examples

Data Zone processing
Inference data is processed within the Microsoft-specified US, EU, or APAC data zone.
Standard pay-per-token
Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.003025
Per month
$30.25

Across 15 published price regions.

Customize workload or price region for Data Zone Standard, No cached input

50% cached input

Per request
$0.00265375
Per month
$26.5375

Across 15 published price regions.

Customize workload or price region for Data Zone Standard, 50% cached input
View 15 region prices for Data Zone Standard

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Data Zone Standard exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Central UScentralus$0.275 $2.20 $0.0275 Not published
East USeastus$0.275 $2.20 $0.0275 Not published
East US 2eastus2$0.275 $2.20 $0.0275 Not published
France Centralfrancecentral$0.275 $2.20 $0.0275 Not published
Germany West Centralgermanywestcentral$0.275 $2.20 $0.0275 Not published
Italy Northitalynorth$0.275 $2.20 $0.0275 Not published
North Central USnorthcentralus$0.275 $2.20 $0.0275 Not published
North Europenortheurope$0.275 $2.20 $0.0275 Not published
Poland Centralpolandcentral$0.275 $2.20 $0.0275 Not published
South Central USsouthcentralus$0.275 $2.20 $0.0275 Not published
Spain Centralspaincentral$0.275 $2.20 $0.0275 Not published
Sweden Centralswedencentral$0.275 $2.20 $0.0275 Not published
West Europewesteurope$0.275 $2.20 $0.0275 Not published
West USwestus$0.275 $2.20 $0.0275 Not published
West US 3westus3$0.275 $2.20 $0.0275 Not published

Global Batch prices and cost examples

Global processing
Inference data may be processed in any Azure region.
Asynchronous Batch
Batch meters cover asynchronous jobs. Microsoft documents Global and Data Zone Batch at a 50% discount, with a 24-hour target for Global Batch; model support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.001375
Per month
$13.75

Across 23 published price regions.

Customize workload or price region for Global Batch, No cached input

50% cached input

Per request
$0.00120625
Per month
$12.0625

Across 23 published price regions.

Customize workload or price region for Global Batch, 50% cached input
View 23 region prices for Global Batch

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Global Batch exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Australia Eastaustraliaeast$0.125 $1.00 $0.0125 Not published
Brazil Southbrazilsouth$0.125 $1.00 $0.0125 Not published
Canada Eastcanadaeast$0.125 $1.00 $0.0125 Not published
East USeastus$0.125 $1.00 $0.0125 Not published
East US 2eastus2$0.125 $1.00 $0.0125 Not published
France Centralfrancecentral$0.125 $1.00 $0.0125 Not published
Germany West Centralgermanywestcentral$0.125 $1.00 $0.0125 Not published
Japan Eastjapaneast$0.125 $1.00 $0.0125 Not published
Korea Centralkoreacentral$0.125 $1.00 $0.0125 Not published
North Central USnorthcentralus$0.125 $1.00 $0.0125 Not published
Norway Eastnorwayeast$0.125 $1.00 $0.0125 Not published
Poland Centralpolandcentral$0.125 $1.00 $0.0125 Not published
South Africa Northsouthafricanorth$0.125 $1.00 $0.0125 Not published
South Central USsouthcentralus$0.125 $1.00 $0.0125 Not published
South Indiasouthindia$0.125 $1.00 $0.0125 Not published
Spain Centralspaincentral$0.125 $1.00 $0.0125 Not published
Sweden Centralswedencentral$0.125 $1.00 $0.0125 Not published
Switzerland Northswitzerlandnorth$0.125 $1.00 $0.0125 Not published
UAE Northuaenorth$0.125 $1.00 $0.0125 Not published
UK Southuksouth$0.125 $1.00 $0.0125 Not published
West Europewesteurope$0.125 $1.00 $0.0125 Not published
West USwestus$0.125 $1.00 $0.0125 Not published
West US 3westus3$0.125 $1.00 $0.0125 Not published

Global Priority prices and cost examples

Global processing
Inference data may be processed in any Azure region.
Priority pay-per-token
Priority processing uses pay-as-you-go token meters for faster responses. Model and deployment support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.00495
Per month
$49.50

Across 25 published price regions.

Customize workload or price region for Global Priority, No cached input

50% cached input

Per request
$0.0043425
Per month
$43.425

Across 25 published price regions.

Customize workload or price region for Global Priority, 50% cached input
View 25 region prices for Global Priority

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Global Priority exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Australia Eastaustraliaeast$0.45 $3.60 $0.045 Not published
Brazil Southbrazilsouth$0.45 $3.60 $0.045 Not published
Canada Eastcanadaeast$0.45 $3.60 $0.045 Not published
Central UScentralus$0.45 $3.60 $0.045 Not published
East USeastus$0.45 $3.60 $0.045 Not published
East US 2eastus2$0.45 $3.60 $0.045 Not published
France Centralfrancecentral$0.45 $3.60 $0.045 Not published
Germany West Centralgermanywestcentral$0.45 $3.60 $0.045 Not published
Japan Eastjapaneast$0.45 $3.60 $0.045 Not published
Korea Centralkoreacentral$0.45 $3.60 $0.045 Not published
North Central USnorthcentralus$0.45 $3.60 $0.045 Not published
North Europenortheurope$0.45 $3.60 $0.045 Not published
Norway Eastnorwayeast$0.45 $3.60 $0.045 Not published
Poland Centralpolandcentral$0.45 $3.60 $0.045 Not published
South Africa Northsouthafricanorth$0.45 $3.60 $0.045 Not published
South Central USsouthcentralus$0.45 $3.60 $0.045 Not published
South Indiasouthindia$0.45 $3.60 $0.045 Not published
Spain Centralspaincentral$0.45 $3.60 $0.045 Not published
Sweden Centralswedencentral$0.45 $3.60 $0.045 Not published
Switzerland Northswitzerlandnorth$0.45 $3.60 $0.045 Not published
UAE Northuaenorth$0.45 $3.60 $0.045 Not published
UK Southuksouth$0.45 $3.60 $0.045 Not published
West Europewesteurope$0.45 $3.60 $0.045 Not published
West USwestus$0.45 $3.60 $0.045 Not published
West US 3westus3$0.45 $3.60 $0.045 Not published

Global Standard prices and cost examples

Global processing
Inference data may be processed in any Azure region.
Standard pay-per-token
Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.00275
Per month
$27.50

Across 27 published price regions.

Customize workload or price region for Global Standard, No cached input

50% cached input

Per request
$0.0024125
Per month
$24.125

Across 27 published price regions.

Customize workload or price region for Global Standard, 50% cached input
View 27 region prices for Global Standard

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-5 mini Global Standard exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Australia Eastaustraliaeast$0.25 $2.00 $0.025 Not published
Brazil Southbrazilsouth$0.25 $2.00 $0.025 Not published
Canada Eastcanadaeast$0.25 $2.00 $0.025 Not published
Central UScentralus$0.25 $2.00 $0.025 Not published
East USeastus$0.25 $2.00 $0.025 Not published
East US 2eastus2$0.25 $2.00 $0.025 Not published
France Centralfrancecentral$0.25 $2.00 $0.025 Not published
Germany West Centralgermanywestcentral$0.25 $2.00 $0.025 Not published
Italy Northitalynorth$0.25 $2.00 $0.025 Not published
Japan Eastjapaneast$0.25 $2.00 $0.025 Not published
Korea Centralkoreacentral$0.25 $2.00 $0.025 Not published
North Central USnorthcentralus$0.25 $2.00 $0.025 Not published
North Europenortheurope$0.25 $2.00 $0.025 Not published
Norway Eastnorwayeast$0.25 $2.00 $0.025 Not published
Poland Centralpolandcentral$0.25 $2.00 $0.025 Not published
South Africa Northsouthafricanorth$0.25 $2.00 $0.025 Not published
South Central USsouthcentralus$0.25 $2.00 $0.025 Not published
South Indiasouthindia$0.25 $2.00 $0.025 Not published
Southeast Asiasoutheastasia$0.25 $2.00 $0.025 Not published
Spain Centralspaincentral$0.25 $2.00 $0.025 Not published
Sweden Centralswedencentral$0.25 $2.00 $0.025 Not published
Switzerland Northswitzerlandnorth$0.25 $2.00 $0.025 Not published
UAE Northuaenorth$0.25 $2.00 $0.025 Not published
UK Southuksouth$0.25 $2.00 $0.025 Not published
West Europewesteurope$0.25 $2.00 $0.025 Not published
West USwestus$0.25 $2.00 $0.025 Not published
West US 3westus3$0.25 $2.00 $0.025 Not published

Model information and limits

Compact GPT model for low-latency assistance and high-volume workloads

Metadata model ID
gpt-5-mini
Metadata match
model
Input modalities
text, image
Output modalities
text
Context window (tokens)
400,000
Maximum input tokens
272,000
Maximum output tokens
128,000
Reasoning
Yes
Tool calling
Yes
Structured output
Not published

Supplemental metadata comes from Models.dev, with verified Azure version corrections where available. Family matches do not establish version-specific limits. Confirm capabilities and lifecycle in the official model catalog.

Price sources and deployment guidance

Rates come from the Azure Retail Prices API. This page covers consumption token meters in Azure Public. Provisioned throughput, training, non-token media meters and sovereign-cloud pricing are separate billing scopes.

Browse all Azure AI model prices