Models and providers

Microsoft Foundry Model Deployment Decision

Select and govern a Microsoft Foundry model deployment using current catalog availability, workload evaluation, identity, region, capacity, cost, and recovery evidence.

Core concept
CurrentNext review due 2026-09-09Content version 0.36.0
Format
Decision Path
Level
Advanced
Audience
Developer, Operator, Leader
Owner
Project42 Editorial
Review cadence
Every 45 days
Prerequisites
An Azure subscription and approved tenant boundary; A defined workload, region, data class, and budget; Permission to inspect catalog availability and quota
01

Start from the actual tenant, region, and resource type

Microsoft Foundry catalog visibility and deployment options depend on resource type, model publisher, region, offer, subscription access, quota, and lifecycle. Confirm the current option in the target tenant rather than assuming a model visible in documentation is deployable everywhere.

Distinguish standard deployments, provisioned throughput, managed compute, and preview access. Record marketplace terms and the identity and network boundary. Treat a preview capability as unavailable for production unless the risk owner explicitly accepts its support and change profile.

02

Evaluate the workload and capacity shape

Use catalog benchmarks to discover candidates, then evaluate with representative application data and tool trajectories. Measure quality, safety, latency, throughput, quota behavior, usage, estimated cost, regional failure, and operational visibility.

Give the deployment a stable application-facing name while recording the exact underlying model and version separately. This supports controlled replacement without silently changing evaluation provenance.

Foundry deployment decision record
text
Task: [WORKLOAD AND DEPLOYMENT DECISION]
Scope: [AZURE TENANT, SUBSCRIPTION, REGION, RESOURCE TYPE, AND DATA CLASS]
Permissions: [DEPLOYER IDENTITY, RBAC, NETWORK, MARKETPLACE, AND SPEND AUTHORITY]
Candidate: [PUBLISHER, EXACT MODEL/VERSION, LIFECYCLE, AND OFFER]
Deployment shape: [STANDARD | PROVISIONED | MANAGED COMPUTE | OTHER]
Capacity: [QUOTA, THROUGHPUT, CONCURRENCY, LATENCY, AND BURST]
Evaluation: [APPLICATION CASES, CATALOG EVIDENCE, THRESHOLDS, AND COST]
Stop conditions: [UNAPPROVED PREVIEW, QUOTA GAP, DATA/REGION GAP, OR FAILED GATE]
Verification: [IDENTITY/NETWORK TEST, LOAD TEST, EVALUATION, AND USER POSTCONDITION]
Recovery: [ROUTE TO VERIFIED DEPLOYMENT, RECONCILE STATE, AND REMOVE FAILED CAPACITY]
03

Expected evidence and verification

Expected evidence includes target tenant and region, resource type, exact catalog item and lifecycle, offer terms, deployment shape, quota, identity and network controls, representative evaluation, load and cost results, stable application route, owner, and expiry.

Verify authentication and authorization from the application identity, network restrictions, request and content behavior, observability, quota failure, failover, and final user result. On failure, stop promotion, restore the verified route, reconcile possible tool writes, and delete or scale down only the failed deployment after evidence is preserved.