Appearance
Foundry observability operations guide
Current consolidated dashboard
The operational dashboard combines native Foundry requests, availability, latency, tokens and billing with gateway, router, MCP and cost-collector views. It uses two-column charts, three compact cost summaries per row and wide detail tables. Completed-migration status is kept in deployment evidence, not a panel. The private overlay identifies the single live dashboard resource and scopes.
Inspect requested versus returned model to understand Azure router selection. First-token time requires streamed content/reasoning/tool output; non-streaming MCP calls have no such measurement. Generation speed is an estimate from stream timing and reported usage. Finish reasons and cached/reasoning token fields are provider-dependent; missing coverage is not zero. Cached/reasoning tokens are subsets, not additions to input/output. Media polling latency is not job duration.
Review HTTP errors, throttles, retries, cancellations, interrupted streams, content-filter error codes, MCP request logs and cost snapshot freshness. Prompt/response content is not retained in gateway telemetry. Cost by meter and forecast come from billing snapshots and can lag usage. Shared platform costs must be labeled separately; a product budget must explicitly define its scopes. Configured action groups are not proof of successful notification delivery. See router operation and gateway behavior.
Scope: Azure AI Foundry
This page describes the Azure AI Foundry target, the hosted-cloud target of ADR-0011. Foundry Local and Azure Local Foundry differ from it in models, features, identity, cost, and operations. Compare all three on Deployment targets.
Daily and weekly review
| Cadence | Review | Evidence | Action |
|---|---|---|---|
| Daily | Budget actual, forecast, anomaly, and workspace quota | Cost Management and alert history | Investigate source before raising any quota or threshold |
| Daily | Failed deployment, delete, Azure incident, or resource health event | Activity Log and active alerts | Establish impact, assign the Foundry owner, and follow the linked runbook |
| Weekly | Foundry inventory, ownership, lifecycle, and expiry | Resource Graph query library | Correct metadata, extend approved lifecycle, or schedule retirement |
| Weekly | Model requests, availability, latency, errors, throttling, and token trend | Foundry dashboard and supported metrics | Compare with an agreed workload baseline before changing an alert |
| Monthly | Diagnostic volume, retention, sampling, and operational value | Workspace usage and Cost Management | Keep, reduce, or remove paid collection |
Alert policy
Every enabled alert needs a named owner, severity, action group, suppression approach, and runbook. Use low-noise scopes and sustained windows. Do not page on latency, token use, or preview project metrics until a workload baseline and response action exist.
For planned maintenance, use an approved alert-processing rule with a narrow scope and time bound. Restore monitoring immediately after the maintenance window and review any suppressed alert history.
Diagnostic approval record
Before enabling Foundry diagnostics, Application Insights, availability tests, or a scheduled query alert, record:
- The exact operational question and source category.
- Data classification, including whether prompt, response, tool, or personal content could be present.
- Reader and writer RBAC, retention, sampling, and access review.
- Daily-volume estimate, expected query or test frequency, monthly cost owner, and review date.
- Alert severity, action group, runbook, and removal condition.
If any item is unknown, leave the optional module disabled.
Foundry-specific response order
- Check Azure Service Health and Resource Health for regional or service impact.
- Check Activity Log for deployment, configuration, identity, or network changes.
- Check supported Foundry metrics for request rate, availability, latency, error, throttling, token, or image-generation evidence.
- Use approved diagnostics or Application Insights only if they were enabled and the data classification permits access.
- Capture the outcome in the workload incident or operations record, then tune the metric threshold or runbook through source control if appropriate.
Model-use investigation
- Select the incident or reporting time window in the model-usage dashboard.
- Start with
ModelRequestssplit byModelDeploymentNameto identify the model deployment and timing. - Compare input, output, and total tokens for the same deployment. This is consumption evidence, not billed currency.
- Review availability, HTTP 429 throttles, and 5xx server errors before assuming a consumer or model problem.
- Use Cost Management for actual and forecast charges. Billing data can arrive later than platform metrics, and currency cost is not inferred from generic token counts.
- Do not enable request, response, trace, or usage diagnostics during an incident without its data-classification and cost approval record.