Lab 11: KServe Inference with HAMi DRA GPU Sharing
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
The same outcome through Kubernetes-native Dynamic Resource Allocation (experimental).
Enforce vGPU count, memory, and compute quotas for HAMi workloads before Pods reach the scheduler.