Skip to content

Model quality and cost review: 2026-09-25 ​

Use small models for routine work, mid-priced capable models for difficult coding and infrastructure, and expensive models for escalations. These are research-based choices, not proven workload winners. The owner subsequently approved applying the starting role map below on September 25. The private overlay records activation and hosted release evidence. The original research itself did not run a paid benchmark.

Azure prices ​

Public USD retail per million tokens, East US global/Global Standard meters, retrieved 2026-09-25 using the Azure Retail Prices API. These are list prices, not deployment spend or negotiated invoice rates. The comparison uses uncached input, short context and non-batch inference. Long context, data zones and caching change rates. Selected raw meter records and reproducible filters are retained in the private overlay.

ModelInput USD/MOutput USD/MSuggested place
GPT-6 Luna0.100.50Routine drafting, extraction and classification candidate
DeepSeek V4 Flash0.190.51Cheap reasoning and coding challenger
Grok 4.1 Fast0.200.50Existing fast route; retain for comparison
GPT-5.6 Luna0.201.20Current router's small tier
Codestral 25010.300.90Narrow code-completion evaluation
Mistral Large 30.501.50Low-cost general-purpose alternative
Kimi K2.7 Code0.954.00Keep coding default pending matched evaluation
Grok 4.31.252.50Another-vendor alternative
Mistral Medium 3.51.507.50Coding challenger, not default promotion
DeepSeek V4 Pro1.743.48Independent reasoning review
Grok 4.62.006.00Primary challenger for deep reasoning, coding and IaC
GPT-6 Sol2.0010.00Difficult coding and IaC candidate
GPT-5.6 Terra2.0012.00Current router's middle tier
GPT-5.6 Sol4.0020.00Current deep role/router high tier
GPT-6 Astra10.0050.00Difficult, consequential escalations

For equal token counts, GPT-6 Sol costs one fifth of Astra and half of GPT-5.6 Sol. Actual savings depend on reasoning/output length, retries, caching and accepted answers. Microsoft also publishes the GPT-6 Azure price table.

Evidence ​

Primary sourceResult and implicationLimitation
OpenAI guidanceLuna for repeatable work, Sol for demanding reasoning/coding, Astra for highest-capability work. Supports a tiered trial.Vendor positioning, not a workload test on these repositories
Kimi reportKimi Code Bench v2 improves from K2.6's 50.9 to 62.0; MCPAtlas 76.0. Credible coding default to retain.Vendor tests with differing harnesses/settings, not proof it beats GPT-6
DeepSeek model cardFlash Max: SWE Verified 79.0, Terminal Bench 2.0 56.9; Pro Max: 80.6 and 67.9. Flash merits a cheap coding trial; Pro merits review work.Max reasoning uses more tokens; results do not transfer to V4.1 or July refreshes
Grok 4.6 reportCursorBench 3.2: 69.9 versus GPT-5.6 Sol's 67.2; DeepSWE 1.1: 65.9 versus 73.0. Results vary by task.Vendor comparisons/settings; no direct proof of best documentation quality
Mistral reportMedium 3.5 SWE-bench Verified 77.6 supports a coding trial.Different harness/token budgets; not directly rankable against DeepSeek's table

These are capability indicators, not a combined leaderboard. We did not find a matched primary-source evaluation of all exact Azure snapshots on PowerShell, Bicep and these documentation tasks. GPT-6's recent release limits operational evidence. A universally best model is not established.

Task recommendations ​

TaskFirst choice to evaluateEscalation or alternative
Simple summaries, formatting, routine documentationGPT-6 LunaDeepSeek V4 Flash or existing Grok 4.1 Fast
Self-contained codeKeep Kimi K2.7 CodeTrial DeepSeek V4 Flash for cheap tasks; GPT-6 Sol for harder work
PowerShell, Bicep, Terraform, architectureCompare GPT-6 Sol and Grok 4.6Astra for unresolved hard tasks or added review
Complex documentation from multiple sourcesGPT-6 Sol or retained Grok 4.6Compare factual correctness and reviewer editing time
Independent design reviewKeep DeepSeek V4 ProGrok 4.6 for another vendor perspective
Adversarial reviewRetain existing Grok reasoning route provisionallyCompare with Grok 4.6 before switching
Automatic selectionKeep explicit Balanced GPT-5.6 subsetEvaluate Cost mode separately; GPT-6 is not documented in this pool

Grok 4.6 is a serious deep reasoning/coding candidate, not only a documentation model. Its output rate is 40 percent below GPT-6 Sol at equal token counts. The published comparisons above are against GPT-5.6 Sol, not GPT-6 Sol; they do not establish which wins on our infrastructure work.

Latest release versus deployed availability ​

Grok 4.20 reasoning and Grok 4.6 are separate generations. Grok 4.6 also reasons; the missing suffix does not mean non-reasoning. Microsoft's Foundry guide documents low/medium/high/xhigh effort for Chat Completions. The private adversary and docs assignments are policy choices, not model limitations or evidence that 4.20 is a better adversarial reviewer. Compare 4.6 for that role.

Grok 4.7 launched September 21 and is newer than 4.6. The deployment snapshot here contains 4.6. A fresh account-scoped East US Foundry catalog query on September 25 returned 4.6 but no 4.7. This is scoped availability evidence, not a claim that 4.7 is unavailable everywhere. Recheck the intended Azure account, SKU, quota and Azure rate before proposing a 4.7 deployment; direct-provider availability/pricing is not Azure availability/pricing.

The owner removed the legacy MCP exclusions for Phi-4 Reasoning, Llama 4 Maverick and the formerly failing DeepSeek V4.1 alias on September 25. All deployed chat/router entries are permitted. The earlier DeepSeek failure was recorded in the retired environment; it did not establish failure in the new environment. Permission is separate from task-quality evaluation, and the two V4.1 deployment aliases remain distinct. This review did not verify a V4.1 Azure meter mapping; do not apply V4 prices to it. Likewise, a Grok 4.2 meter label does not prove the rate of every 4.20 deployment variant. Media, audio and embeddings require modality-specific evaluation and are outside this chat-role ranking.

Promotion criteria ​

Adopted starting role map ​

For a concrete cost-aware starting configuration: fast, cheap-bulk and routine docs use GPT-6 Luna; code stays Kimi K2.7 Code; deep and adversary use Grok 4.6; iac uses GPT-6 Sol; second-opinion stays DeepSeek V4 Pro; auto keeps the Balanced GPT-5.6 subset. Keep fast as the omitted-role default. The owner approved this configuration for activation after the research review. Workload-quality evaluation remains necessary; operational smoke tests verify connectivity and role resolution, not which model is best at a task. Use explicit stronger-role calls for complex documentation. Astra is an escalation option rather than a default. The current MCP does not implement automatic quality-based escalation chains; callers make those decisions.

Evaluate subsequent changes ​

Use fixed representative tasks, exact versions and equal acceptance criteria. Measure accepted answers, syntax/tests, unsupported claims, reviewer editing time, p50/p95 latency, billed input/output including reasoning, retries and cost per accepted result. Start with a small capped evaluation. Record results before changing roles.

Check API compatibility: OpenAI guidance requires Responses for Astra tool calling, and for Sol/Luna reasoning with tools. The current MCP adapter delegates text via Chat Completions without giving the model its caller's tools. Direct agent integration needs separate compatibility checks; the gateway does not translate APIs. Follow model evaluation and router policy synchronization.