Skip to content

Project Practices ​

Project Cases ​

DocumentDescription
Multi-Compute Pool Heterogeneous Inference Scheduling Best PracticeBest practice for heterogeneous inference scheduling across multiple compute pools
Single-Node Multi-Card Multi-Model Deployment Best PracticeBest practice for partitioning one multi-card NPU/GPU node into reusable specs and validating multi-model deployment
Model Auto-Download and Inference Template Validation Best PracticeOperator model download and inference-template setup, followed by an End User creating an instance, publishing a private model, and validating it in Playground and with cURL