Skip to content

Observability & Troubleshooting ​

This scenario helps operators, providers, and callers decide whether a problem belongs to calls, model services, cloud deployments, On-Prem resources, or metering before opening the corresponding logs and monitoring views.

Applicable Roles ​

  • End User, Model Provider, and Platform Operator investigating issues within their permitted scope

Target Outcome ​

  • The issue includes role, time range, model or resource ID, and a reproducible symptom.
  • The issue is narrowed to call, deployment, job, node, device, or metering level.
  • Logs, events, monitoring, and usage use the same time range.
  • Evidence is redacted and sufficient for the next owner.

Before You Start ​

  1. Record time, account role, tenant, subsystem, and page entry.
  2. Record a redacted model, deployment, instance, job, or request identifier.
  3. Classify the symptom as visibility, creation, runtime, call, performance, or usage.
  4. Do not send full prompts, responses, tokens, keys, or internal endpoints in tickets or chats.

Routing Table ​

SymptomFirst EntryManual
Model API call fails or response is abnormalMy Calls or Customer Call LogsCall Logs, Customer Call Logs
Success rate, latency, or token use is abnormalCall AnalyticsMy Call Analytics, Customer Analytics
Cloud deployment fails or is unreachableDeployment details, events, and monitoringMy Deployments
On-Prem job is pending or failedJob monitoring, instance events, and logsJob Monitoring, Instances
Node or accelerator is abnormalNode Statistics and Device MonitoringNode Statistics, Device Monitoring
Quota, usage, or amount is abnormalQuota, metering details, and model usageOn-Prem Metering & Monitoring, Model Usage & Earnings

General Sequence ​

  1. Reproduce and capture the first error instead of only the final cascading error.
  2. Confirm account, tenant, region, model, and time filters.
  3. Move from user-visible state to events and logs, then to node or device monitoring.
  4. For call issues, compare request logs with model service state.
  5. For resource issues, compare job state, node capacity, and device health.
  6. For usage issues, validate runtime records before metering details and period summaries.

Use the On-Prem Monitoring Overview during layer identification to compare cluster, node, device, and workload signals in the same time range.

Compare monitoring signals in the same time range

Completion Checklist ​

Purpose: These are the exit criteria for the current feature task. Use them to decide whether the result is observable and reviewable and whether you can continue to the next step in the scenario. They do not repeat the procedure; if any item fails, follow the troubleshooting section below.

CheckPass Criteria
1The issue layer and current owner are clear.
2Error, log, event, and monitoring times align.
3Impact is classified as one request, instance, tenant, or the platform.
4The same conditions were retested after mitigation or repair.
5Handoff includes entry, steps, expected and actual results, time, and redacted evidence.

Troubleshooting and Common Mistakes ​

  • Reviewing only aggregate monitoring without failed events or request logs.
  • Using different time, region, or tenant filters across views.
  • Treating permission-driven invisibility as missing resources.
  • Mixing quota, account credit, and cluster capacity failures.
  • Copying complete requests or credentials into evidence.