Skip to content

Azure Local SCOM monitoring research

This is the raw research inventory for the local Management Pack. It intentionally contains more signals than the default product. Collectable does not mean enabled, health-impacting, or pageable.

The current development implementation is documentation- and contract-validated. Command shape, provider availability, counter names, event behavior, latency, recovery, and overhead still require the representative Azure Local and SCOM lab.

Research spikes

SpikeQuestionRequired output
Support matrixWhich Azure Local, SCOM, node-count, stretched-cluster, and hardware combinations are supported?Version matrix and tested fixture IDs
Local topologyDo stable keys reconcile across node contributions, owner change, restart, replacement, and grooming?Object/relationship snapshots and lifecycle timing
Health ServiceWhich current fault types, severities, associations, and recovery timings appear in each target release?Fault catalog, injected condition, observed state, recovery
StorageWhich pool, volume, CSV, disk, storage job, enclosure, cache, and repair signals are safe defaults?Raw inventory, workflow mapping, cardinality, cost
Network ATCWhich intent properties and status values occur during deployment, convergence, drift, remediation, and failure?State machine, events, fault injections
Registration/platformWhich local registration, Arc, MOC, resource-bridge, and extension states are stable and supported?Local source contract and Microsoft-support boundaries
LifecycleWhich solution update environment, update, run, step, and health states are actionable?Update state machine and maintenance behavior
PerformanceWhich local counters and Cluster Performance History series exist by version/hardware?Counter export, unit/instance mapping, collection cost
Events and logsWhich providers, channels, IDs, schemas, and correlation keys are stable?Exported manifests, captured events, suppression keys
Threshold engineeringWhat duration, recovery, reserve, and topology context makes each candidate actionable?Evidence worksheet and tuned profile proposal
SCOM runtimeDo embedded providers cook down and run under every target HealthService runtime?Workflow traces, event log, agent CPU/memory
Release lifecycleDo import, upgrade, override preservation, rollback, and removal behave safely?Repeatable certification report

Authoritative local state sources

AreaCandidate interfacesData
ClusterGet-Cluster, Get-ClusterNode, Get-ClusterQuorum, Get-ClusterGroup, Get-ClusterResourceIdentity, membership, votes, quorum/witness, roles, owners, state
CSVGet-ClusterSharedVolumeOwner, path, online state, redirected access
Health ServiceGet-HealthFaultFault ID/type, severity, reason, recommendation, object, physical location, create/update/remove
Health actionsGet-HealthAction and storage jobs where availableRepair/remediation action, progress, duration, failure
Storage subsystemGet-StorageSubSystem, Get-StoragePool, Get-VirtualDisk, Get-VolumeHealth, operational state, capacity, resiliency, read-only, allocation
Physical storageGet-PhysicalDisk, Get-StorageReliabilityCounterIdentity, serial, media/bus/usage, location, health, wear, temperature, errors
Network ATCGet-NetIntent, Get-NetIntentStatusIntent, traffic role, adapters, configuration/provisioning state, last update
RegistrationGet-AzureStackHCIRegistration, connection, last connected, Azure resource name/URI, verification state
LifecycleGet-SolutionUpdateEnvironment, Get-SolutionUpdate, Get-SolutionUpdateRunCurrent solution, package, readiness, health, progress, failed steps
Platform servicesGet-Service plus product-specific diagnosticsCluster, Arc, extension, MOC, VM-management service presence/state
HostCIM and Windows countersHardware, OS, CPU, memory, network, disk
PipelineSCOM workflow state and product eventsLast success, exception, duration, object count, freshness

Microsoft documents that Health Service faults provide severity, reason, recommended action, affected object, and physical location, and that root-cause analysis suppresses consequential noise. That is why the default design pages on the root fault instead of every derived drive symptom: Health Service faults.

Performance metric inventory

Microsoft currently publishes more than 60 Azure Local platform metrics through the telemetry and diagnostics extension. The cloud list is also the cross-track parity checklist for local SCOM research; each metric must be mapped to Cluster Performance History, a stable Windows counter, a scripted provider, or an explicit unsupported result.

Node and compute

  • Percentage CPU; Percentage CPU Guest; Percentage CPU Host.
  • Cluster node Memory Total, Available, and Used.
  • Percentage Memory, Percentage Memory Guest, and Percentage Memory Host.
  • Cluster node CSV cache Read Hit, Read Hit rate, and Read Miss.
  • Cluster node Storage Degraded.
  • Optional GPU: Percentage GPU, Percentage GPU Memory, GPU Temperature, GPU Graphics Clock Speed, and GPU Memory Clock Speed when supported GPU partitioning and drivers are present.

Physical drives

  • Read Operations/sec, Write Operations/sec, and combined Read and Write Operations/sec.
  • Read Bytes/sec, Write Bytes/sec, and combined Read and Write throughput.
  • Read latency, write latency, and average latency.
  • Total capacity and used capacity.
  • Reliability candidates: temperature, wear, read/write/uncorrectable errors, and power-on hours, subject to vendor exposure and lab validation.

Network adapters

  • Network In/sec, Network Out/sec, and Network Total/sec.
  • RDMA inbound, outbound, and total bandwidth.
  • Candidate diagnostics: link speed/state, errors/discards, RDMA operational state, SMB Direct connections, SET membership, and intent-to-adapter mapping.

Volumes and CSVs

  • Read Operations/sec, Write Operations/sec, and combined operations.
  • Read Bytes/sec, Write Bytes/sec, and combined throughput.
  • Read latency, write latency, and average latency.
  • Total size and available size.
  • Candidate state: health, operational status, owner, redirected access/reason, resiliency, repair progress, and free-capacity trend.

VHD and VM series

The Azure platform exposes VHD operations, throughput, latency, current/max size and VM CPU, memory, pressure, and virtual-network throughput. They remain raw research inputs only. The core Azure Local MP monitors infrastructure, not customer guest workloads; these series are candidates for a future workload companion or collection-only dashboard, not default infrastructure health.

See the current Microsoft metric names, units, aggregations, and dimensions: Monitor Azure Local with Azure Monitor Metrics.

Cluster Performance History provides curated live cluster, server, and volume metrics through one cmdlet. Microsoft documents those values as point-in-time when queried from PowerShell: Cluster Performance History.

Events and logs

The default development MP uses four established System-log Failover Clustering conditions:

ProviderEventMeaning
Microsoft-Windows-FailoverClustering1135Node removed from active cluster membership
Microsoft-Windows-FailoverClustering1069 / 1205Cluster resource or role failure
Microsoft-Windows-FailoverClustering5120CSV access failure
Microsoft-Windows-FailoverClustering5142CSV no longer accessible

The lab spike must enumerate every available event channel and provider manifest on each target version before adding more event IDs. Candidate families include Failover Clustering operational and diagnostic channels, Storage Spaces Direct/Space Manager, Storage Health, Network ATC, Azure Local registration, Arc agents, MOC, resource bridge, solution updates, and SCOM HealthService. Event absence is never proof of Healthy; state-bearing signals require a current-state source.

What should be monitored first

PriorityDefault decisionExamples
MustHealth-impacting and normally alertableCluster Service, node membership, quorum, root Health Service faults, unhealthy pool/volume, CSV inaccessible, failed Network ATC intent, registration failure, required platform service failure, failed update, pipeline failure
ShouldHealth-impacting but alert policy needs evidenceSingle physical-disk degradation, repair duration, available capacity trend, Arc disconnection duration, update readiness warnings
CouldUseful diagnostics or optional collectionEnclosure detail, cache efficiency, storage jobs, RDMA bandwidth, GPU metrics, detailed reliability counters
Collection onlyTime series without default health/pageCPU, memory, throughput, IOPS, latency, host/guest split
Excluded from coreDifferent ownership boundaryGuest OS, application, customer VM expected state, pod/deployment health, cloud ARM resource configuration

Threshold and duration questions

A threshold is not accepted because it is common or round. Each proposal must record source, topology, unit, aggregation, sample interval, consecutive samples or duration, reset threshold, maintenance behavior, alert severity, auto-resolution, missing-data state, and lab evidence.

  • Memory: use available host memory, host/guest split, pressure, paging, failover reserve, and trend. Do not page solely because used memory reached 75 percent.
  • CPU: distinguish host from guest demand, require sustained duration, and verify schedulability and workload impact.
  • Volume: combine percentage and absolute free bytes for large and small volumes; validate growth rate and repair reserve.
  • Latency: segment by media, operation, volume/disk, cache state, queue depth, and sustained duration.
  • Physical disks: map the deployment fault-tolerance model and active repair before escalating a count.
  • Registration: separate brief connectivity loss from sustained disconnection and distinguish local platform service failure from Azure service availability.
  • Updates: available content is informational; failed preparation/install and blocked critical health checks are actionable.

Current implementation boundary

The authored development baseline includes the Must state families, 12 conservative Windows performance collections, the four event rules above, a diagnostic task, operator knowledge, DA rollup, and override generation. Four high-cardinality or counter-availability-sensitive collections begin disabled. No release default is called certified until the research spikes and lab matrix produce evidence.

Released under the MIT License.