Skip to content

Devices ​

Feature Overview ​

ItemContent
Applicable RoleOperator
Navigation PathAI Infra(On-Prem) > Monitoring > Devices
Page Route/powerone/monitor/device
Managed ObjectConfiguration, status, and relationships on Devices

Beginner Explanation ​

Device monitoring is like an accelerator dashboard. It observes GPU/NPU utilization, VRAM, temperature, and health status to determine whether compute cards can continue hosting tasks.

Terms ​

TermDescription
Device UtilizationCurrent compute utilization of GPU/NPU.
VRAM UsageAccelerator VRAM occupation ratio.
TemperatureDevice operating temperature.

Confirm prerequisites for Accelerator devices such as GPU/NPU, VRAM, utilization, temperature, and health status, follow Main Operations, run Result Validation, and continue to the next page.

First-Time User Notes ​

Confirm that the task involves Configuration, status, and relationships on Devices, and then follow the recommended order. If fields or state differ from expectations, check prerequisites before continuing downstream.

Prerequisites ​

  1. The current account has device monitoring view permissions.
  2. The target cluster has deployed device plugins and can report GPU/NPU metrics.
  3. Device model, VRAM, temperature, and health status can be collected.
  4. The accelerator model or job scope to focus on has been confirmed.

Page Description ​

Use this page to inspect GPU and NPU utilization, device memory, temperature, and health state.

Devices

Device monitoring is used to view GPU/NPU utilization, VRAM, temperature, and health status. Operators can use it to determine whether accelerators are offline, overheating, out of VRAM, or occupied by a single task for a long time.

The following figure shows the device monitoring page.

Main Operations ​

Filter and View Device Monitoring ​

  1. Go to AI Infrastructure > On-Prem > Monitoring > Device Monitoring.
  2. Confirm the region and resource pool in the upper-right corner, and filter by cluster, node, device type, device status, or time range.
  3. View the device list and overall running status, and verify device ID, device type, node, cluster, region/AZ, and device status.
  4. Review accelerator usage, VRAM usage, temperature, health status, bound jobs, and exception information to identify unavailable devices, insufficient VRAM, or hardware exceptions.

View device monitoring

Investigate Abnormal Device Metrics ​

  1. When an accelerator is offline, overheated, running out of VRAM, or throwing driver errors, note the device ID and node.
  2. Keeping the same time range, navigate to Job Monitoring to confirm if a large model training or high-concurrency inference job is occupying the card.
  3. Switch to Node Statistics to check the host node load, network, and driver service health.
  4. If a hardware failure or PCIe bus error is verified, coordinate with hardware engineers for isolation and maintenance rather than force-killing workloads without review.

Key Focus ​

  • Whether devices are identified and continuously reported.
  • Whether VRAM and utilization are close to limits.
  • Whether temperature, error counts, or health status are abnormal.

Parameter Quick Reference ​

Field NameRequiredField TypeExampleDescription
Device IDSystem-generatedTextGPU-0Distinguishes multiple devices on the same node.
Device TypeYesTextNVIDIA A800Shows GPU/NPU or other accelerator type and model.
NodeConditionally requiredTextnode-gpu-01Locates the node where the device resides.
ClusterConditionally requiredTextcluster-prod-aLocates the cluster to which the device belongs.
Region / AZConditionally requiredDrop-downWuhan / AZ ALimits the resource location to which the device belongs.
Device StatusSystem-generatedStatusNormalShows whether the device is available, alerted, or offline.
Accelerator UsageSystem-generatedPercentage92%Determines whether compute units are under high load.
VRAM UsageSystem-generatedPercentage / Capacity62 GB / 80 GBDetermines whether a model or job occupies all VRAM.
TemperatureSystem-generatedNumber71°CHelps judge cooling and hardware health.
Health StatusSystem-generatedStatusNormalShows whether the device is available, alerted, or offline.
Bound JobSystem-generatedText / NumberRunning jobShows jobs currently associated with or occupying the device.
Time RangeConditionally requiredDate rangeLast 1 hourControls the query window for statistic cards, trend charts, and list data.

Pitfalls ​

  • Full VRAM does not necessarily mean compute is fully loaded. Judge together with utilization.
  • Temperature exceptions should be escalated to operations promptly for hardware and cooling checks.
  • When devices are invisible, check drivers, plugins, and node status first.
  • Device monitoring may have collection latency. Do not judge hardware faults based only on a single instant metric.
  • Device exceptions should be investigated together with node status, job status, scheduling events, device plugins, and node logs.
  • High VRAM watermark does not necessarily mean a device fault. Judge together with bound jobs and model specifications.
  • Do not write real device IDs, node names, node IPs, cluster IDs, resource pool IDs, tenant information, internal metric keys, or test data in the document.

Configuration Rules and Impact ​

  • VRAM watermark directly affects model startup: When VRAM is insufficient, instance creation may fail even if total cluster resources look sufficient.
  • View temperature and health together: High temperature, missing cards, or driver exceptions can all cause job failures.
  • Device dimension is suitable for hotspot location: When cluster watermarks are normal but jobs are slow, use the device dimension to confirm whether a single-card hotspot exists.
  • Model differences affect schedulability: The same specification may require a specific GPU/NPU model, driver, or compute capability.

Result Validation ​

Check ItemSuccess SignalIf Abnormal
Page loadDevices charts or lists are visibleCheck monitoring permission and whether collection is available in the selected region
ScopeTime range, region, and object count match the investigationClear filters and restore them one at a time to avoid mixed scopes
FreshnessUpdate time is within the expected collection intervalCheck collection interval, connection, and alerts in system or monitoring configuration
CorrelationAn abnormal metric can be linked to a cluster, node, device, or jobKeep the same time range and cross-check adjacent monitoring pages and object details

FAQ ​

No Data on Devices ​

Symptom:

The page opens, but charts or lists are empty.

Possible Causes:

  • No job ran in the selected time.
  • collection is unavailable in the region.
  • the role lacks metric permission.

Solution:

  1. Expand the time range and reset filters
  2. verify regional monitoring capability
  3. compare an adjacent monitoring page.

Devices Is Not Updating ​

Symptom:

The data does not change for an extended period.

Possible Causes:

  • The next collection cycle has not arrived.
  • the collector is abnormal.
  • the page is cached.

Solution:

  1. Check update time
  2. inspect collector status and alerts
  3. refresh with the same time range.

Devices Differs from Adjacent Pages ​

Symptom:

The same object has different values on two monitoring pages.

Possible Causes:

  • Aggregation granularity differs.
  • time range or time zone differs.
  • filters target different objects.

Solution:

  1. Align time range and time zone
  2. verify aggregation scope
  3. clear and restore filters one at a time.

Symptom:

The metric or details entry does not lead to the expected object.

Possible Causes:

  • The object ended or was removed.
  • the role cannot see it.
  • relationship identifiers differ.

Solution:

  1. Record object and time
  2. check its list state
  3. ask the Operator to verify visibility.

A Spike Cannot Be Reproduced ​

Symptom:

A spike was recorded, but current details are normal.

Possible Causes:

  • The spike was brief.
  • sampling is coarse.
  • the job has ended.

Solution:

  1. Lock the spike interval
  2. compare job and node events
  3. retain a sanitized screenshot and object identifier.

Notes ​

  • Device serial numbers, node locations, and internal hardware IDs should be sanitized.
  • Empty utilization does not necessarily indicate an exception. Combine it with the task time range.
  • Device health exceptions should be handled according to hardware procedures first.
  • Before device fault judgment, cross-check with node status, job status, scheduling events, device plugins, and node logs.
  • Documentation examples must not include real device IDs, node names, node IPs, cluster IDs, resource pool IDs, tenant information, internal metric keys, or test data.

Next Steps ​

  1. When VRAM is high, enter job monitoring to locate occupying tasks.
  2. When temperature or health is abnormal, contact operations to handle hardware or drivers.
  3. When model resources are insufficient, review accelerator configuration and specification association.