Azure Speed Test

Prices by deployment type

All rates are USD per 1 million tokens. Each row uses its own deployment type and published price regions. A range reflects regional differences within that type. Missing prices mean no published meter in this snapshot; they do not mean free usage. A retail price does not guarantee deployment availability or quota.

GPT-4o mini 0718 OpenAI prices in USD per 1 million tokens.
Deployment typeInput / 1MOutput / 1MCached input / 1MCache write / 1MPrice regions
Data Zone Batch $0.083 $0.33 Not published Not published 16
Data Zone Standard $0.165 $0.66 $0.083 Not published 17
Global Batch $0.075 $0.30 Not published Not published 27
Global Standard $0.15 $0.60 $0.075 Not published 28
Regional Standard $0.165 - $0.198 $0.66 - $0.792 $0.083 - $0.099 Not published 11

The tables contain only this OpenAI listing. Other providers, versions and similarly named Azure products can have different meters. Some token categories may be missing in individual regions; expand the regional tables for exact coverage.

How token costs are calculated

For each region and deployment type, multiply each token category by its matching rate, then divide the sum by 1,000,000. Monthly estimates multiply that request cost by the number of requests. Ranges use complete workloads priced in each region, so input and output prices from different regions are never combined.

Request cost = (uncached input tokens × input rate + cached input tokens × cached input rate + cache-write input tokens × cache-write rate + output tokens × output rate) / 1,000,000.

The examples below use zero cache writes. The cached example assumes half of the input tokens are eligible cache reads on every request; actual cache hit rates and any cache creation charges depend on the deployment. Missing cache rates are not replaced with ordinary input rates. Estimates cover these token meters only, before taxes, negotiated discounts and any separate hosting, tools, media or provisioned-capacity charges.

Some model token limits are not published in the metadata. These are arithmetic cost examples; verify that your deployment supports the workload before using it.

Data Zone Batch prices and cost examples

Data Zone processing
Inference data is processed within the Microsoft-specified US, EU, or APAC data zone.
Asynchronous Batch
Batch meters cover asynchronous jobs. Microsoft documents Global and Data Zone Batch at a 50% discount, with a 24-hour target for Global Batch; model support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.000579
Per month
$5.79

Across 16 published price regions.

Customize workload or price region for Data Zone Batch, No cached input

50% cached input

Estimate unavailable

Published cached input prices are missing for this workload in one or more matching price regions. Select a region with complete prices.

Customize workload or price region for Data Zone Batch, 50% cached input
View 16 region prices for Data Zone Batch

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-4o mini 0718 Data Zone Batch exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Central UScentralus$0.083 $0.33 Not published Not published
East USeastus$0.083 $0.33 Not published Not published
East US 2eastus2$0.083 $0.33 Not published Not published
France Centralfrancecentral$0.083 $0.33 Not published Not published
Germany West Centralgermanywestcentral$0.083 $0.33 Not published Not published
North Central USnorthcentralus$0.083 $0.33 Not published Not published
North Europenortheurope$0.083 $0.33 Not published Not published
Poland Centralpolandcentral$0.083 $0.33 Not published Not published
South Central USsouthcentralus$0.083 $0.33 Not published Not published
Southeast Asiasoutheastasia$0.083 $0.33 Not published Not published
Spain Centralspaincentral$0.083 $0.33 Not published Not published
Sweden Centralswedencentral$0.083 $0.33 Not published Not published
West Europewesteurope$0.083 $0.33 Not published Not published
West USwestus$0.083 $0.33 Not published Not published
West US 2westus2$0.083 $0.33 Not published Not published
West US 3westus3$0.083 $0.33 Not published Not published

Data Zone Standard prices and cost examples

Data Zone processing
Inference data is processed within the Microsoft-specified US, EU, or APAC data zone.
Standard pay-per-token
Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.001155
Per month
$11.55

Across 17 published price regions.

Customize workload or price region for Data Zone Standard, No cached input

50% cached input

Per request
$0.001032
Per month
$10.32

Across 17 published price regions.

Customize workload or price region for Data Zone Standard, 50% cached input
View 17 region prices for Data Zone Standard

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-4o mini 0718 Data Zone Standard exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Central UScentralus$0.165 $0.66 $0.083 Not published
East USeastus$0.165 $0.66 $0.083 Not published
East US 2eastus2$0.165 $0.66 $0.083 Not published
France Centralfrancecentral$0.165 $0.66 $0.083 Not published
Germany West Centralgermanywestcentral$0.165 $0.66 $0.083 Not published
Italy Northitalynorth$0.165 $0.66 $0.083 Not published
North Central USnorthcentralus$0.165 $0.66 $0.083 Not published
North Europenortheurope$0.165 $0.66 $0.083 Not published
Poland Centralpolandcentral$0.165 $0.66 $0.083 Not published
South Central USsouthcentralus$0.165 $0.66 $0.083 Not published
Southeast Asiasoutheastasia$0.165 $0.66 $0.083 Not published
Spain Centralspaincentral$0.165 $0.66 $0.083 Not published
Sweden Centralswedencentral$0.165 $0.66 $0.083 Not published
West Europewesteurope$0.165 $0.66 $0.083 Not published
West USwestus$0.165 $0.66 $0.083 Not published
West US 2westus2$0.165 $0.66 $0.083 Not published
West US 3westus3$0.165 $0.66 $0.083 Not published

Global Batch prices and cost examples

Global processing
Inference data may be processed in any Azure region.
Asynchronous Batch
Batch meters cover asynchronous jobs. Microsoft documents Global and Data Zone Batch at a 50% discount, with a 24-hour target for Global Batch; model support varies.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.000525
Per month
$5.25

Across 27 published price regions.

Customize workload or price region for Global Batch, No cached input

50% cached input

Estimate unavailable

Published cached input prices are missing for this workload in one or more matching price regions. Select a region with complete prices.

Customize workload or price region for Global Batch, 50% cached input
View 27 region prices for Global Batch

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-4o mini 0718 Global Batch exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Australia Eastaustraliaeast$0.075 $0.30 Not published Not published
Brazil Southbrazilsouth$0.075 $0.30 Not published Not published
Canada Eastcanadaeast$0.075 $0.30 Not published Not published
Central UScentralus$0.075 $0.30 Not published Not published
East USeastus$0.075 $0.30 Not published Not published
East US 2eastus2$0.075 $0.30 Not published Not published
France Centralfrancecentral$0.075 $0.30 Not published Not published
Germany West Centralgermanywestcentral$0.075 $0.30 Not published Not published
Italy Northitalynorth$0.075 $0.30 Not published Not published
Japan Eastjapaneast$0.075 $0.30 Not published Not published
Korea Centralkoreacentral$0.075 $0.30 Not published Not published
North Central USnorthcentralus$0.075 $0.30 Not published Not published
North Europenortheurope$0.075 $0.30 Not published Not published
Norway Eastnorwayeast$0.075 $0.30 Not published Not published
Poland Centralpolandcentral$0.075 $0.30 Not published Not published
South Africa Northsouthafricanorth$0.075 $0.30 Not published Not published
South Central USsouthcentralus$0.075 $0.30 Not published Not published
South Indiasouthindia$0.075 $0.30 Not published Not published
Southeast Asiasoutheastasia$0.075 $0.30 Not published Not published
Spain Centralspaincentral$0.075 $0.30 Not published Not published
Sweden Centralswedencentral$0.075 $0.30 Not published Not published
Switzerland Northswitzerlandnorth$0.075 $0.30 Not published Not published
UAE Northuaenorth$0.075 $0.30 Not published Not published
UK Southuksouth$0.075 $0.30 Not published Not published
West Europewesteurope$0.075 $0.30 Not published Not published
West USwestus$0.075 $0.30 Not published Not published
West US 3westus3$0.075 $0.30 Not published Not published

Global Standard prices and cost examples

Global processing
Inference data may be processed in any Azure region.
Standard pay-per-token
Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.00105
Per month
$10.50

Across 28 published price regions.

Customize workload or price region for Global Standard, No cached input

50% cached input

Per request
$0.0009375
Per month
$9.375

Across 28 published price regions.

Customize workload or price region for Global Standard, 50% cached input
View 28 region prices for Global Standard

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-4o mini 0718 Global Standard exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
Australia Eastaustraliaeast$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Brazil Southbrazilsouth$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Canada Eastcanadaeast$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Central UScentralus$0.15 $0.60 $0.075 Not published
East USeastus$0.15 $0.60 $0.075 Not published
input and output:
cached input:
East US 2eastus2$0.15 $0.60 $0.075 Not published
input and output:
cached input:
France Centralfrancecentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Germany West Centralgermanywestcentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Italy Northitalynorth$0.15 $0.60 $0.075 Not published
Japan Eastjapaneast$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Korea Centralkoreacentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
North Central USnorthcentralus$0.15 $0.60 $0.075 Not published
input and output:
cached input:
North Europenortheurope$0.15 $0.60 $0.075 Not published
Norway Eastnorwayeast$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Poland Centralpolandcentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
South Africa Northsouthafricanorth$0.15 $0.60 $0.075 Not published
input and output:
cached input:
South Central USsouthcentralus$0.15 $0.60 $0.075 Not published
input and output:
cached input:
South Indiasouthindia$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Southeast Asiasoutheastasia$0.15 $0.60 $0.075 Not published
Spain Centralspaincentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Sweden Centralswedencentral$0.15 $0.60 $0.075 Not published
input and output:
cached input:
Switzerland Northswitzerlandnorth$0.15 $0.60 $0.075 Not published
input and output:
cached input:
UAE Northuaenorth$0.15 $0.60 $0.075 Not published
input and output:
cached input:
UK Southuksouth$0.15 $0.60 $0.075 Not published
input and output:
cached input:
West Europewesteurope$0.15 $0.60 $0.075 Not published
input and output:
cached input:
West USwestus$0.15 $0.60 $0.075 Not published
input and output:
cached input:
West US 2westus2$0.15 $0.60 $0.075 Not published
West US 3westus3$0.15 $0.60 $0.075 Not published
input and output:
cached input:

Regional Standard prices and cost examples

Single-region processing
Inference data is processed in the deployment region. Microsoft names the pay-per-token deployment type Standard; this directory uses Regional Standard to make its scope explicit.
Standard pay-per-token
Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.

Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.

No cached input

Per request
$0.001155 - $0.001386
Per month
$11.55 - $13.86

Across 11 published price regions.

Customize workload or price region for Regional Standard, No cached input

50% cached input

Per request
$0.001032 - $0.0012375
Per month
$10.32 - $12.375

Across 11 published price regions.

Customize workload or price region for Regional Standard, 50% cached input
View 11 region prices for Regional Standard

Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.

GPT-4o mini 0718 Regional Standard exact regional rates in USD per 1 million tokens.
Azure Public price regionInput / 1MOutput / 1MCached input / 1MCache write / 1MMeter effective dates (UTC)
East USeastus$0.165 $0.66 $0.083 Not published
input and output:
cached input:
East US 2eastus2$0.165 $0.66 $0.083 Not published
input and output:
cached input:
North Central USnorthcentralus$0.165 $0.66 $0.083 Not published
input and output:
cached input:
North Europenortheurope$0.198 $0.792 $0.099 Not published
South Central USsouthcentralus$0.165 $0.66 $0.083 Not published
input and output:
cached input:
Southeast Asiasoutheastasia$0.198 $0.792 $0.099 Not published
Sweden Centralswedencentral$0.165 $0.66 $0.091 Not published
input and output:
cached input:
West Central USwestcentralus$0.165 $0.66 $0.083 Not published
West USwestus$0.165 $0.66 $0.083 Not published
input and output:
cached input:
West US 2westus2$0.165 $0.66 $0.083 Not published
West US 3westus3$0.165 $0.66 $0.083 Not published
input and output:
cached input:

Model information and limits

Small omni GPT for cheap multimodal assistance and production-scale traffic

Metadata model ID
gpt-4o-mini
Metadata match
family
Input modalities
Not published
Output modalities
Not published
Context window (tokens)
128,000
Maximum input tokens
Not published
Maximum output tokens
16,384
Reasoning
Not published
Tool calling
Yes
Structured output
Not published

Supplemental metadata comes from Models.dev, with verified Azure version corrections where available. Family matches do not establish version-specific limits. Confirm capabilities and lifecycle in the official model catalog.

Price sources and deployment guidance

Rates come from the Azure Retail Prices API. This page covers consumption token meters in Azure Public. Provisioned throughput, training, non-token media meters and sovereign-cloud pricing are separate billing scopes.

Browse all Azure AI model prices