Phi-3 Small 8K Instruct Azure Pricing
Microsoft prices for Azure Phi Models in Microsoft Foundry. Published rates cover 1 deployment profile across 10 Azure Public price regions.
Prices by deployment type
All rates are USD per 1 million tokens. Each row uses its own deployment type and published price regions. A range reflects regional differences within that type. Missing prices mean no published meter in this snapshot; they do not mean free usage. A retail price does not guarantee deployment availability or quota.
| Deployment type | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Price regions |
|---|---|---|---|---|---|
| Regional Standard | $0.15 | $0.60 | Not published | Not published | 10 |
The tables contain only this Microsoft listing. Other providers, versions and similarly named Azure products can have different meters. Some token categories may be missing in individual regions; expand the regional tables for exact coverage.
How token costs are calculated
For each region and deployment type, multiply each token category by its matching rate, then divide the sum by 1,000,000. Monthly estimates multiply that request cost by the number of requests. Ranges use complete workloads priced in each region, so input and output prices from different regions are never combined.
Request cost = (uncached input tokens × input rate + cached input tokens × cached input rate + cache-write input tokens × cache-write rate + output tokens × output rate) / 1,000,000.
The examples below use zero cache writes. The cached example assumes half of the input tokens are eligible cache reads on every request; actual cache hit rates and any cache creation charges depend on the deployment. Missing cache rates are not replaced with ordinary input rates. Estimates cover these token meters only, before taxes, negotiated discounts and any separate hosting, tools, media or provisioned-capacity charges.
Some model token limits are not published in the metadata. These are arithmetic cost examples; verify that your deployment supports the workload before using it.
Regional Standard prices and cost examples
- Single-region processing
- Inference data is processed in the deployment region. Microsoft names the pay-per-token deployment type Standard; this directory uses Regional Standard to make its scope explicit.
- Standard pay-per-token
- Input and output usage is billed by token. Standard deployments provide best-effort throughput; reserved PTU pricing is separate.
Example workload: 3,000 input and 1,000 output tokens per request, with 10,000 requests per month.
No cached input
- Per request
- $0.00105
- Per month
- $10.50
Across 10 published price regions.
Customize workload or price region for Regional Standard, No cached input50% cached input
Estimate unavailable
Published cached input prices are missing for this workload in one or more matching price regions. Select a region with complete prices.
Customize workload or price region for Regional Standard, 50% cached inputView 10 region prices for Regional Standard
Price region identifies the Azure retail meter; processing location follows the deployment type described above. Dates show the latest effective meter in each row, not when prices were last checked. All token rates are USD per 1 million tokens.
| Azure Public price region | Input / 1M | Output / 1M | Cached input / 1M | Cache write / 1M | Meter effective dates (UTC) |
|---|---|---|---|---|---|
| Central UScentralus | $0.15 | $0.60 | Not published | Not published | |
| East USeastus | $0.15 | $0.60 | Not published | Not published | |
| East US 2eastus2 | $0.15 | $0.60 | Not published | Not published | |
| North Central USnorthcentralus | $0.15 | $0.60 | Not published | Not published | |
| South Central USsouthcentralus | $0.15 | $0.60 | Not published | Not published | |
| Sweden Centralswedencentral | $0.15 | $0.60 | Not published | Not published | |
| West Central USwestcentralus | $0.15 | $0.60 | Not published | Not published | |
| West USwestus | $0.15 | $0.60 | Not published | Not published | |
| West US 2westus2 | $0.15 | $0.60 | Not published | Not published | |
| West US 3westus3 | $0.15 | $0.60 | Not published | Not published |
Model information and limits
No exact supplemental model metadata is available for this Azure listing. Input, output and context limits are unknown. The Azure retail prices above remain separate from capability and deployment-availability information.
Price sources and deployment guidance
Rates come from the Azure Retail Prices API. This page covers consumption token meters in Azure Public. Provisioned throughput, training, non-token media meters and sovereign-cloud pricing are separate billing scopes.
- Microsoft Foundry pricing
Review the official platform pricing overview and Azure purchasing guidance.
- Microsoft Foundry model catalog
Browse the broader official model catalog, providers, APIs, and capabilities.
- Foundry Models sold by Azure
Check model IDs, context windows, APIs, capabilities, and lifecycle guidance.
- Azure OpenAI pricing
Review Microsoft's public model tables and current Azure OpenAI list prices.
- Foundry deployment types
Understand data processing scope, pay-per-token, Batch, and provisioned PTUs.
- Azure Retail Prices API
See the Microsoft API documentation for the retail meter data used here.