Direct Answer: Deploying open-source LLMs on Red Hat OpenShift AI replaces manual Docker scripts with Kubernetes-native autoscaling and automated hardware discovery. Pair KServe's SingleModelServing architecture with the vLLM serving runtime and the NVIDIA GPU Operator to achieve continuous batching, PagedAttention memory optimization, and enterprise-grade TLS ingress.
Start here
Intended reader: Platform engineers, MLOps specialists, and enterprise architects hosting LLMs on Kubernetes. Practical outcome: A fully configured, reproducible OpenShift AI deployment serving 8B–14B parameter models with vLLM and KServe. As enterprise data sovereignty mandates move models into private clouds, OpenShift AI provides a declarative control plane that automates GPU driver lifecycle, vLLM continuous batching, and secure API ingress.
Prerequisites: Node Feature Discovery and the NVIDIA GPU Operator
Before provisioning serving pods, OpenShift worker nodes must expose GPU hardware as schedulable Kubernetes resources. Install Node Feature Discovery (NFD) to scan PCIe buses and label worker nodes with GPU family tags. Follow with the NVIDIA GPU Operator, which automatically compiles and deploys NVIDIA driver containers, container toolkits, and DCGM monitoring pods without manual host configuration.
Serving runtime: KServe with vLLM PagedAttention
OpenShift AI supports two serving paradigms: ModelMesh (for dense multi-model CPU workloads) and SingleModelServing via KServe (the standard for LLMs). Within your Data Science Project, select the vLLM Serving Runtime template. vLLM implements PagedAttention, which partitions KV cache memory into non-contiguous blocks, virtually eliminating memory fragmentation and enabling continuous batching across concurrent client queries.
Cluster configuration: storage mount, resources and TLS route
Connect an S3-compatible object store (such as Ceph, AWS S3, or MinIO) containing model weights in SafeTensors format. Allocate dedicated GPU slices in the ServingRuntime manifest:
| Deployment Parameter | Recommended Specification | Operational Impact |
|---|---|---|
| Resource Limits | nvidia.com/gpu: 1, memory: 32Gi | Guarantees dedicated VRAM allocation for 8B models without pod eviction |
| Serving Engine | vLLM v0.6+ Container Image | Delivers 2.5x–4x higher throughput via continuous request scheduling |
| Model Weight Source | S3 Bucket / Internal MinIO | Enables rapid pod cold starts and zero-downtime rolling weight rollouts |
| Ingress Protocol | OpenShift Route (Edge TLS) | Terminates HTTPS at router layer and exposes standard OpenAI-compatible API |
Production checklist: day-two operations and autoscaling
Verify these four operational checkpoints before routing customer traffic to your inference cluster:
- Configure Horizontal Pod Autoscaler (HPA) targeting Prometheus metric 'vllm:num_requests_waiting'.
- Enable dedicated node taints (e.g. 'nvidia.com/gpu=present:NoSchedule') to prevent non-AI workloads from scheduling on GPU nodes.
- Set '--gpu-memory-utilization 0.90' in vLLM args to reserve 10% VRAM for peak context activation surges.
- Verify readiness probes are mapped to the '/health' endpoint to prevent traffic routing during model weight loading.
Try this next
Evaluating local vs edge model performance? Compare quantization tradeoffs in our guide on Reduce Local LLM VRAM Usage.
Sources & further reading
Primary sources checked Sep 22, 2026. Vendor statements are attributed; editorial advice is our own.
- 1
- 2KServe vLLM Runtime Guidelines & API Specifications ↗KServe Project
- 3
Help us keep this useful. Send a correction or a primary source →



