Building a sovereign LLM inference platform: LiteLLM, vLLM and NIM on H200 behind an F5

One private, OpenAI-compatible endpoint now serves every internal AI application in our data centre, running entirely on our own H200 GPUs. This post walks through how I built it: Kubernetes on bare metal, the NVIDIA operators, vLLM and NIM serving the models, LiteLLM as the gateway, and an F5 load balancer on the uplink.

The goal: one sovereign endpoint

The scenario is an energy enterprise where firewall analysis, ticket automation and call summarisation all need LLMs, and none of that data may leave the building. Public APIs were off the table. Giving each team its own GPU server would have wasted the hardware and multiplied the operational work.

So the design goal was simple: teams call one URL with one API key format, and the platform decides which model answers, on which GPU, and how much they may spend.

The hardware is six GPU worker nodes, each with 8 NVIDIA H200 NVL cards. That is 1,128 GB of GPU memory per server, 48 GPUs and 6,768 GB in total. Three CPU-only servers run the Kubernetes control plane. Everything is Dell.

By the end of this post you will have seen every layer, in the order I built it:

  1. The physical topology and what each network carries
  2. Kubernetes (RKE2) on an immutable OS, with a control plane and GPU workers
  3. The GPU Operator and Network Operator
  4. Cluster networking with Cilium, Istio and MetalLB
  5. Secrets and GitOps
  6. A shared model cache
  7. Model servers: vLLM and NIM
  8. LiteLLM as the gateway
  9. The F5 load balancer in front of it all
  10. Validation and the mistakes I would avoid

A note on names: hostnames, IP ranges and tenant details in this post are placeholders. The design and the commands are real.

The topology: three networks, one way in

The platform runs on three separate networks, and the F5 sits on top of the uplink as the only way in.

NetworkCarriesWhere it lives
Uplink and F5Client requests and the Kubernetes APIThe F5 above the uplink
GPU fabricInbound requests and GPU-to-GPU RDMA traffic2 spines, 2 leaves, 2 RDMA NICs per node
ManagementNode management NICs and iDRACDell management switch, 3 CPU servers

Keeping these apart is deliberate. A saturated GPU fabric cannot stop you reaching a node’s iDRAC, and a management-network problem cannot slow a model down.

Step 1: Kubernetes on bare metal with RKE2

I split the cluster into three CPU-only control plane nodes and six GPU workers, and I kept AI workloads off the control plane entirely. The control plane runs etcd and the API server. The GPU workers run nothing except the things that need GPUs.

RoleCountGPUsTaintWhat runs there
Control plane3noneCriticalAddonsOnlyetcd, API server, scheduler
GPU worker68× H200 NVL eachnvidia.com/gpudrivers, model servers

All nine machines run an immutable, transactional OS (SLE Micro). The root filesystem is read-only, which is great for consistency and means you cannot install things the usual way. Two consequences shaped the build:

  • RKE2 installs from its tarball into /opt/rke2, not through a package manager.
  • Anything that needs kernel modules or drivers has to arrive through containers, not zypper install.

Bootstrap the first server

Before starting the service, write the config. The tls-san entries matter: they must include the load balancer name you will publish for the API, or clients will reject the certificate later.

# /etc/rancher/rke2/config.yaml (first control plane node)
token: <shared-cluster-secret>
tls-san:
  - <k8s-api-fqdn>
  - <k8s-api-vip>
cni: cilium
disable-kube-proxy: true
node-taint:
  - "CriticalAddonsOnly=true:NoExecute"


curl -sfL https://get.rke2.io | INSTALL_RKE2_METHOD=tar sh -
systemctl enable --now rke2-server.service
export KUBECONFIG=/etc/rancher/rke2/rke2.yaml
/var/lib/rancher/rke2/bin/kubectl get nodes

The other two control plane nodes get the same file plus a server: line pointing at the first node, then start the same service. Three etcd members give you quorum with one node down.

Join the GPU workers

Workers use the agent role. I label and taint them at join time so nothing lands on a GPU node by accident.

# /etc/rancher/rke2/config.yaml (GPU worker)
server: https://<k8s-api-fqdn>:9345
token: <shared-cluster-secret>
node-label:
  - "node-role.kubernetes.io/gpu-worker=true"
node-taint:
  - "nvidia.com/gpu=present:NoSchedule"

curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=agent INSTALL_RKE2_METHOD=tar sh -
systemctl enable --now rke2-agent.service

The supervisor port 9345 is used only while joining. The API server itself is 6443. Whatever load balancer you put in front of the control plane has to forward both, which comes back in the F5 section.

All nine nodes should now show Ready. The GPU workers have no GPU capability yet, because Kubernetes does not know the cards exist. That is the next step.

Step 2: Teaching Kubernetes about the GPUs and the fabric

Two NVIDIA operators turn a plain Kubernetes node into a GPU node: the GPU Operator handles the cards, the Network Operator handles the RDMA NICs. Neither touches the host OS, which is exactly what an immutable root filesystem needs.

GPU Operator

The GPU Operator deploys a set of DaemonSets on every GPU node: the driver, the container toolkit, the device plugin, GPU Feature Discovery (which labels nodes with GPU model and memory) and the DCGM exporter for metrics. On RKE2 the container toolkit must be told where RKE2 keeps its containerd config, otherwise pods never see the NVIDIA runtime.

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia && helm repo update

helm install gpu-operator nvidia/gpu-operator \
  -n gpu-operator --create-namespace \
  --set toolkit.env[0].name=CONTAINERD_CONFIG \
  --set toolkit.env[0].value=/var/lib/rancher/rke2/agent/etc/containerd/config.toml.tmpl \
  --set toolkit.env[1].name=CONTAINERD_SOCKET \
  --set toolkit.env[1].value=/run/k3s/containerd/containerd.sock \
  --set toolkit.env[2].name=CONTAINERD_RUNTIME_CLASS \
  --set toolkit.env[2].value=nvidia \
  --set toolkit.env[3].name=CONTAINERD_SET_AS_DEFAULT \
  --set-string toolkit.env[3].value=true

The operator’s DaemonSets already tolerate the nvidia.com/gpu taint from step 1, so they land on the GPU workers and nowhere else. Give the driver container a few minutes to build and load, then check what each node advertises:

kubectl -n gpu-operator get pods -o wide
kubectl describe node <gpu-node> | grep -A2 nvidia.com/gpu
# Capacity:  nvidia.com/gpu: 8

Eight per node, 48 across the cluster. If a node shows 0, the driver pod is almost always the culprit, so read its logs first.

Network Operator

The Network Operator does for the RDMA NICs what the GPU Operator does for the GPUs. It installs the NIC driver stack and an RDMA device plugin, so pods can request an RDMA resource the same way they request nvidia.com/gpu. Everything is described in one custom resource:

apiVersion: mellanox.com/v1alpha1
kind: NicClusterPolicy
metadata:
  name: nic-cluster-policy
spec:
  ofedDriver:
    image: doca-driver
    repository: nvcr.io/nvidia/mellanox
    version: <doca-driver-version>
  rdmaSharedDevicePlugin:
    image: k8s-rdma-shared-dev-plugin
    repository: ghcr.io/mellanox
    version: <plugin-version>
    config: |
      {"configList": [{
        "resourceName": "rdma_shared_device_a",
        "rdmaHcaMax": 63,
        "selectors": {"ifNames": ["<rdma-nic-1>", "<rdma-nic-2>"]}
      }]}

Field names and image versions move between releases, so copy the current sample for your operator version rather than trusting this snippet blindly.

One honest point about when this matters. A model that fits inside one node uses NVLink and PCIe between its GPUs, and never touches the RDMA fabric. The two-NIC leaf-and-spine fabric earns its keep for multi-node serving and distributed training. Install it early anyway, because retrofitting a fabric under running workloads is far more painful than doing it up front.

Step 3: Cluster networking, three tools with three jobs

I used Cilium for pod networking, Istio for HTTP-level control, and MetalLB to give services real IP addresses the F5 can target. Each does one job, and keeping those jobs separate is what made the platform debuggable.

LayerToolJob
L3-L4Cilium (eBPF)Pod networking, network policy, kube-proxy replacement
L7IstiomTLS between services, retries, request-level policy, ingress gateway
Service IPsMetalLBHands out routable IPs to LoadBalancer services

Cilium

Because step 1 set cni: cilium and disable-kube-proxy: true, Cilium must take over service load balancing. RKE2 lets you tune its bundled chart with a HelmChartConfig dropped into the manifests directory on a control plane node:

Cilium

Because step 1 set cni: cilium and disable-kube-proxy: true, Cilium must take over service load balancing. RKE2 lets you tune its bundled chart with a HelmChartConfig dropped into the manifests directory on a control plane node:

# /var/lib/rancher/rke2/server/manifests/rke2-cilium-config.yaml
apiVersion: helm.cattle.io/v1
kind: HelmChartConfig
metadata:
  name: rke2-cilium
  namespace: kube-system
spec:
  valuesContent: |-
    kubeProxyReplacement: true
    k8sServiceHost: <k8s-api-fqdn>
    k8sServicePort: 6443

Check it with cilium status and confirm every node reports the agent as ready. A cluster with no working CNI leaves every pod in ContainerCreating, so this is the first thing to look at when nothing starts.

MetalLB

On bare metal, a Service of type LoadBalancer stays <pending> forever unless something answers for it. MetalLB is that something. I gave it a small dedicated range on the server network and advertised it in layer 2 mode.

helm repo add metallb https://metallb.github.io/metallb
helm install metallb metallb/metallb -n metallb-system --create-namespace
helm repo add metallb https://metallb.github.io/metallb
helm install metallb metallb/metallb -n metallb-system --create-namespace
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: ingress-pool
  namespace: metallb-system
spec:
  addresses:
    - <ingress-vip-range>
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: ingress-l2
  namespace: metallb-system
spec:
  ipAddressPools:
    - ingress-pool
helm repo add metallb https://metallb.github.io/metallb
helm install metallb metallb/metallb -n metallb-system --create-namespace
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: ingress-pool
  namespace: metallb-system
spec:
  addresses:
    - <ingress-vip-range>
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: ingress-l2
  namespace: metallb-system
spec:
  ipAddressPools:
    - ingress-pool

Layer 2 mode is the simplest and works well when the F5 and the cluster share a VLAN. If your design needs the VIPs routed across subnets, use BGP mode instead and peer with the leaf switches.

Istio

Istio gives the platform mTLS between services and a single ingress gateway that the F5 talks to. I installed the three charts in order:

helm repo add istio https://istio-release.storage.googleapis.com/charts
helm install istio-base istio/base -n istio-system --create-namespace
helm install istiod istio/istiod -n istio-system --wait
helm install istio-ingress istio/gateway -n istio-system

The gateway Service is type LoadBalancer, so MetalLB assigns it an address from the pool. That address is the pool member the F5 will point at later.

One gotcha specific to LLM traffic: responses stream token by token and can run for minutes. Check the idle and stream timeouts on every hop, including Istio routes and the F5 profile, because the shortest one silently cuts long generations.

Step 4: Secrets and GitOps, so nothing is configured by hand

Every component after this point is deployed by Argo CD from Git, and every secret comes from a vault instead of a manifest. That rule kept the platform reproducible: any node or namespace can be rebuilt from the repository plus the vault.

External Secrets

The platform needs real secrets: the gateway master key, the database connection string, a Hugging Face token and an NGC API key. None of them belong in Git. The External Secrets Operator reads them from Azure Key Vault and turns them into ordinary Kubernetes Secrets.

helm repo add external-secrets https://charts.external-secrets.io
helm install external-secrets external-secrets/external-secrets \
  -n external-secrets --create-namespace --set installCRDs=true

One store object tells the operator how to authenticate to the vault. This is the single secret you create by hand, the credential of an Entra ID application that can read the vault, and it lives only in the cluster.

helm repo add metallb https://metallb.github.io/metallb
helm install metallb metallb/metallb -n metallb-system --create-namespace
apiVersion: metallb.io/v1beta1
kind: IPAddressPool
metadata:
  name: ingress-pool
  namespace: metallb-system
spec:
  addresses:
    - <ingress-vip-range>
---
apiVersion: metallb.io/v1beta1
kind: L2Advertisement
metadata:
  name: ingress-l2
  namespace: metallb-system
spec:
  ipAddressPools:
    - ingress-pool

Layer 2 mode is the simplest and works well when the F5 and the cluster share a VLAN. If your design needs the VIPs routed across subnets, use BGP mode instead and peer with the leaf switches.

Istio

Istio gives the platform mTLS between services and a single ingress gateway that the F5 talks to. I installed the three charts in order:

helm repo add istio https://istio-release.storage.googleapis.com/charts
helm install istio-base istio/base -n istio-system --create-namespace
helm install istiod istio/istiod -n istio-system --wait
helm install istio-ingress istio/gateway -n istio-system

The gateway Service is type LoadBalancer, so MetalLB assigns it an address from the pool. That address is the pool member the F5 will point at later.

One gotcha specific to LLM traffic: responses stream token by token and can run for minutes. Check the idle and stream timeouts on every hop, including Istio routes and the F5 profile, because the shortest one silently cuts long generations.

Step 4: Secrets and GitOps, so nothing is configured by hand

Every component after this point is deployed by Argo CD from Git, and every secret comes from a vault instead of a manifest. That rule kept the platform reproducible: any node or namespace can be rebuilt from the repository plus the vault.

External Secrets

The platform needs real secrets: the gateway master key, the database connection string, a Hugging Face token and an NGC API key. None of them belong in Git. The External Secrets Operator reads them from Azure Key Vault and turns them into ordinary Kubernetes Secrets.

helm repo add external-secrets https://charts.external-secrets.io
helm install external-secrets external-secrets/external-secrets \
  -n external-secrets --create-namespace --set installCRDs=true

One store object tells the operator how to authenticate to the vault. This is the single secret you create by hand, the credential of an Entra ID application that can read the vault, and it lives only in the cluster.

apiVersion: external-secrets.io/v1beta1
kind: ClusterSecretStore
metadata:
  name: azure-kv
spec:
  provider:
    azurekv:
      tenantId: <tenant-id>
      vaultUrl: https://<vault-name>.vault.azure.net
      authSecretRef:
        clientId:
          name: azure-kv-credentials
          namespace: external-secrets
          key: client-id
        clientSecret:
          name: azure-kv-credentials
          namespace: external-secrets
          key: client-secret

Then each namespace declares which vault entries it wants:

apiVersion: external-secrets.io/v1beta1
kind: ExternalSecret
metadata:
  name: litellm-secrets
  namespace: litellm
spec:
  refreshInterval: 1h
  secretStoreRef:
    name: azure-kv
    kind: ClusterSecretStore
  target:
    name: litellm-secrets
  data:
    - secretKey: LITELLM_MASTER_KEY
      remoteRef: {key: litellm-master-key}
    - secretKey: DATABASE_URL
      remoteRef: {key: litellm-database-url}
    - secretKey: HF_TOKEN
      remoteRef: {key: huggingface-token}
    - secretKey: NGC_API_KEY
      remoteRef: {key: ngc-api-key}

Rotating a secret is now a vault operation. The operator picks up the change at the next refresh, and the pods restart on their own schedule.

Argo CD

Our Argo CD instance runs in Azure, not in this cluster. I registered the GPU cluster with it, which is why the Kubernetes API has to be reachable through the F5 on port 6443.

argocd cluster add <kube-context> --name gpu-cluster

Each layer of this post is one Argo CD Application. Ordering matters, because a custom resource cannot exist before the operator that defines its CRD. Sync waves handle that.

WaveWhat syncs
0Cilium config, MetalLB, External Secrets
1Istio, GPU Operator, Network Operator
2Cluster policies, IP pools, secret stores
3Model cache, model servers
4LiteLLM
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: litellm
  namespace: argocd
  annotations:
    argocd.argoproj.io/sync-wave: "4"
spec:
  project: ai-platform
  source:
    repoURL: <git-repo-url>
    path: platform/litellm
    targetRevision: main
  destination:
    name: gpu-cluster
    namespace: litellm
  syncPolicy:
    automated: {prune: true, selfHeal: true}
    syncOptions: [CreateNamespace=true]

With selfHeal on, a manual kubectl edit in production is reverted within minutes. That can feel restrictive at first, and it is also what keeps Git and the cluster honest with each other.

Step 5: A shared model cache, so weights are downloaded once

All model weights live on one shared NFS volume that every GPU node mounts, so a model is downloaded once and read many times. Without it, six nodes each pull the same hundreds of gigabytes through the corporate proxy, and every pod restart repeats the download.

The arithmetic explains why this matters. A model’s weights take roughly its parameter count times the bytes per parameter. A 122B-parameter model needs about 122 GB at 8-bit precision and about 244 GB at 16-bit. A 2 TB volume holds a handful of large models plus their variants, and nothing more.

The volume

I exposed the NAS as a ReadWriteMany volume. The mount options matter for large sequential reads, and the Retain policy protects the weights if someone deletes the claim by mistake.

apiVersion: v1
kind: PersistentVolume
metadata:
  name: model-cache
spec:
  capacity:
    storage: 2Ti
  accessModes: [ReadWriteMany]
  persistentVolumeReclaimPolicy: Retain
  mountOptions: [nfsvers=4.1, hard, nconnect=8]
  nfs:
    server: <nfs-server>
    path: /model-cache
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: model-cache
  namespace: nim-service
spec:
  accessModes: [ReadWriteMany]
  storageClassName: ""
  volumeName: model-cache
  resources:
    requests:
      storage: 2Ti

Downloading through the proxy

The cluster has no direct internet access. Hugging Face Hub and NVIDIA NGC are reached through the corporate web proxy. I run the download as a Kubernetes Job so it uses the cluster network, the vault secrets and the shared volume, and leaves a record in Git.

apiVersion: batch/v1
kind: Job
metadata:
  name: fetch-model
  namespace: nim-service
spec:
  backoffLimit: 2
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: fetch
          image: python:3.12-slim
          env:
            - {name: HTTPS_PROXY, value: "http://<proxy-host>:<port>"}
            - {name: NO_PROXY, value: ".svc,.cluster.local,<pod-cidr>"}
            - name: HF_TOKEN
              valueFrom: {secretKeyRef: {name: litellm-secrets, key: HF_TOKEN}}
          command: ["sh", "-c"]
          args:
            - pip install -q "huggingface_hub" &&
              hf download <org>/<model> --local-dir /models/<model-dir>
          volumeMounts:
            - {name: cache, mountPath: /models}
      volumes:
        - name: cache
          persistentVolumeClaim: {claimName: model-cache}

The Job reads the token from a Secret in its own namespace, so the ExternalSecret from step 4 must also exist in nim-service.

Two things that cost time here:

  • If your proxy inspects TLS, it re-signs every certificate with a corporate root. Mount that CA into the Job and set SSL_CERT_FILE and REQUESTS_CA_BUNDLE, or every download fails with a certificate error that looks like a network problem.
  • NIM containers run as a non-root user. The NFS export must allow that user to write, or NIM cannot populate its cache and exits at startup.

Before moving on, confirm the weights are visible from a GPU node by running a throwaway pod that mounts the claim and lists the directory.

Step 6: Serving the models with vLLM and NIM

I serve Hugging Face models with vLLM and NVIDIA-catalogue models with NIM, and both expose the same OpenAI-compatible API. That common API is the reason the gateway in the next step can treat them as interchangeable backends.

AliasModelEnginePurpose
frontierGLM-5.2vLLMLargest general model, hardest tasks
advancedQwen3.5-122BNIMHigh-quality general and vision-capable work
standardQwen3.6-35BvLLMEveryday workloads, cheapest per token
embeddingsbge-m3vLLMVector embeddings for retrieval
rerankgte-rerankervLLMRe-ranking retrieved passages
speechqwen3-ttsvLLMText to speech

The alias is what applications see. The model behind an alias can change without a single application edit, which is the whole point of the gateway.

Placement rules I followed

GPU placement decides both performance and how much of the cluster one model can starve.

  • Keep a model’s tensor-parallel group inside one node. The 8 GPUs in a node talk over NVLink and PCIe, which is far faster than crossing the network.
  • A model that needs all 8 GPUs owns a node. Smaller models share nodes, and one GPU is enough for embeddings and reranking.
  • Never mix a large model and a small one on the same GPU. Memory fragmentation and noisy neighbours make latency unpredictable.

A vLLM deployment

This is the shape of every vLLM model server. Only the model path, the served name and the parallelism change between models.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-standard
  namespace: <models-namespace>
spec:
  replicas: 1
  selector:
    matchLabels: {app: vllm-standard}
  template:
    metadata:
      labels: {app: vllm-standard}
    spec:
      nodeSelector:
        node-role.kubernetes.io/gpu-worker: "true"
      tolerations:
        - {key: nvidia.com/gpu, operator: Exists, effect: NoSchedule}
      containers:
        - name: vllm
          image: vllm/vllm-openai:<version>
          args:
            - --model=/models/<model-dir>
            - --served-model-name=standard
            - --tensor-parallel-size=4
            - --gpu-memory-utilization=0.90
            - --max-model-len=<context-length>
          ports:
            - containerPort: 8000
          resources:
            limits:
              nvidia.com/gpu: 4
          readinessProbe:
            httpGet: {path: /health, port: 8000}
            initialDelaySeconds: 60
            periodSeconds: 10
          volumeMounts:
            - {name: cache, mountPath: /models, readOnly: true}
            - {name: shm, mountPath: /dev/shm}
      volumes:
        - name: cache
          persistentVolumeClaim: {claimName: model-cache}
        - name: shm
          emptyDir: {medium: Memory, sizeLimit: 16Gi}
---
apiVersion: v1
kind: Service
metadata:
  name: vllm-standard
  namespace: <models-namespace>
spec:
  selector: {app: vllm-standard}
  ports:
    - {port: 8000, targetPort: 8000}

Three details that decide whether it starts:

  • The /dev/shm volume is not optional. Multi-GPU inference passes tensors through shared memory, and the container default is far too small.
  • The GPU limit and --tensor-parallel-size must match, or vLLM refuses to start.
  • The readiness probe needs a long initial delay. Loading a large model from NFS takes minutes, and a too-eager probe sends traffic to a pod that is still loading.

A NIM deployment

NIM packages the model and an optimised inference engine into one container image from NVIDIA’s registry. It needs two credentials: an image pull secret for the registry, and the NGC key to fetch model files on first start. Point its cache at the shared volume so the model is fetched once.

      containers:
        - name: nim
          image: nvcr.io/nim/<org>/<model>:<tag>
          env:
            - name: NGC_API_KEY
              valueFrom: {secretKeyRef: {name: litellm-secrets, key: NGC_API_KEY}}
            - {name: NIM_CACHE_PATH, value: /opt/nim/.cache}
            - {name: NIM_SERVED_MODEL_NAME, value: advanced}
          ports:
            - containerPort: 8000
          resources:
            limits:
              nvidia.com/gpu: 8
          readinessProbe:
            httpGet: {path: /v1/health/ready, port: 8000}
            initialDelaySeconds: 120
            periodSeconds: 15
          volumeMounts:
            - {name: cache, mountPath: /opt/nim/.cache}
      imagePullSecrets:
        - name: ngc-registry

NVIDIA also ships a NIM Operator that manages caches and services as custom resources. I would evaluate it once the number of NIM models grows, but a plain Deployment is easier to reason about for the first few.

When the pods are Ready, test each one directly from inside the cluster before the gateway exists:

kubectl -n <models-namespace> run curl --rm -it --image=curlimages/curl -- \
  curl -s http://vllm-standard:8000/v1/models

If that returns the served model name, the model layer works and everything above it is plumbing.

Step 7: LiteLLM, the single front door for every model

LiteLLM turns six different model servers into one API, and adds the things raw model servers lack: per-team keys, budgets, rate limits, retries and usage tracking. Applications talk to LiteLLM using the OpenAI client library they already know, and never learn where a model runs.

What LiteLLM needs

  • A PostgreSQL database for keys, teams, spend and logs. I used a managed PostgreSQL service outside the cluster, so the gateway stays stateless and the state survives a cluster rebuild. Azure requires sslmode=require in the connection string.
  • The master key and database URL, delivered by the ExternalSecret from step 4.
  • A config.yaml that maps each public alias to a backend service.

The configuration

Each model_name is what applications request. Each litellm_params block points at the in-cluster Service of a model server. Because vLLM and NIM both speak the OpenAI protocol, LiteLLM treats them all as openai/ backends.

model_list:
  - model_name: frontier
    litellm_params:
      model: openai/frontier
      api_base: http://vllm-frontier.<models-namespace>.svc.cluster.local:8000/v1
      api_key: none
  - model_name: advanced
    litellm_params:
      model: openai/advanced
      api_base: http://nim-advanced.<models-namespace>.svc.cluster.local:8000/v1
      api_key: none
  - model_name: standard
    litellm_params:
      model: openai/standard
      api_base: http://vllm-standard.<models-namespace>.svc.cluster.local:8000/v1
      api_key: none
  - model_name: bge-m3
    litellm_params:
      model: openai/bge-m3
      api_base: http://vllm-embed.<models-namespace>.svc.cluster.local:8000/v1
      api_key: none

router_settings:
  num_retries: 2
  timeout: 600
  fallbacks:
    - frontier: [advanced]

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

litellm_settings:
  drop_params: true

Two lines earn their place. The fallbacks entry means that if the largest model is down or overloaded, requests quietly go to the next tier instead of failing. The timeout of 600 seconds keeps long generations from being cut off by the gateway.

The deployment

Run at least three replicas behind one Service, since the gateway is stateless.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: litellm
  namespace: litellm
spec:
  replicas: 3
  selector:
    matchLabels: {app: litellm}
  template:
    metadata:
      labels: {app: litellm}
    spec:
      containers:
        - name: litellm
          image: ghcr.io/berriai/litellm:<version>
          args: ["--config", "/app/config.yaml", "--port", "4000"]
          envFrom:
            - secretRef: {name: litellm-secrets}
          ports:
            - containerPort: 4000
          readinessProbe:
            httpGet: {path: /health/readiness, port: 4000}
          volumeMounts:
            - {name: config, mountPath: /app/config.yaml, subPath: config.yaml}
      volumes:
        - name: config
          configMap: {name: litellm-config}
---
apiVersion: v1
kind: Service
metadata:
  name: litellm
  namespace: litellm
spec:
  selector: {app: litellm}
  ports:
    - {port: 4000, targetPort: 4000}

With several replicas, per-key rate limits are only accurate if the replicas share a counter. That is what a Redis instance is for in LiteLLM, and I recommend adding one before you enforce strict limits.

Exposing it through Istio

The Istio ingress gateway from step 3 receives the traffic and routes it to the Service.

apiVersion: networking.istio.io/v1
kind: Gateway
metadata:
  name: llm-gateway
  namespace: litellm
spec:
  selector:
    istio: ingress
  servers:
    - port: {number: 80, name: http, protocol: HTTP}
      hosts: ["<llm-fqdn>"]
---
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: litellm
  namespace: litellm
spec:
  hosts: ["<llm-fqdn>"]
  gateways: [llm-gateway]
  http:
    - route:
        - destination: {host: litellm, port: {number: 4000}}
      timeout: 0s

The timeout: 0s disables Istio’s route timeout so streamed answers are not cut off. TLS is handled one hop earlier, on the F5.

Keys, teams and budgets

Every application gets its own virtual key, restricted to the aliases it needs and limited in rate. That is what stops one runaway script from taking the whole platform down.

curl -X POST https://<llm-fqdn>/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "key_alias": "firewall-analysis",
    "models": ["standard", "bge-m3"],
    "tpm_limit": 200000,
    "rpm_limit": 60,
    "metadata": {"team": "network-security"}
  }'

For people, I sign administrators in to the LiteLLM admin UI with our identity provider (Microsoft Entra ID) instead of a shared password. LiteLLM reads the Entra application details from environment variables, so they come from the vault like everything else. Check LiteLLM’s current licensing for how many SSO users the free tier allows before you plan around it.

Step 8: The F5 on top of the uplink

The F5 is the only entry point into the platform: it owns the public name of the LLM endpoint and the name of the Kubernetes API, and everything behind it stays on private addresses. Clients reach the F5, the F5 sends traffic down the uplink into the fabric, and the Istio gateway takes it from there.

Two virtual servers do all the work:

Virtual serverPortPool membersHealth check
LLM endpoint443Istio ingress VIP (from MetalLB)HTTP GET on the gateway readiness path
Kubernetes API6443 and 9345The three control plane nodesTCP connect

The second row is the reason step 1 asked you to put the load balancer name in tls-san, and the reason Argo CD in Azure can manage this cluster.

The LLM virtual server

I terminate TLS on the F5 with the enterprise certificate, so applications trust the endpoint with no extra configuration. If your security policy demands encryption all the way in, add a server-side SSL profile and enable TLS on the Istio gateway as well.

# Health monitor: the Host header must match the Istio VirtualService host
tmsh create ltm monitor http litellm-health defaults-from http \
  send "GET /health/readiness HTTP/1.1\r\nHost: <llm-fqdn>\r\nConnection: close\r\n\r\n" \
  recv "200"

tmsh create ltm pool llm-istio-pool \
  members add { <istio-ingress-vip>:80 } monitor litellm-health

# Long idle timeout so streamed generations are not cut off
tmsh create ltm profile tcp llm-tcp defaults-from tcp idle-timeout 600

tmsh create ltm profile client-ssl llm-clientssl defaults-from clientssl \
  cert <cert-name> key <key-name>

tmsh create ltm virtual llm-vs \
  destination <f5-vip>:443 ip-protocol tcp pool llm-istio-pool \
  profiles add { llm-tcp http llm-clientssl } \
  source-address-translation { type automap }

Three settings prevented the failures that are hardest to diagnose:

  • The monitor sends the real Host header. Istio routes by hostname, so a monitor without it gets a 404 and the F5 marks a healthy pool as down.
  • The TCP profile idle timeout is 600 seconds. The F5 default is shorter than a long generation, and the symptom is an answer that stops mid-sentence.
  • Do not enable HTTP compression or response buffering on this virtual server. Streaming depends on tokens flowing through as they are produced.

SNAT automap makes the return traffic come back through the F5. It is the simplest choice while the F5 and the MetalLB range share a VLAN.

The Kubernetes API virtual server

This one is plain TCP, because Kubernetes clients and the API server negotiate TLS themselves.

tmsh create ltm pool k8s-api-pool monitor tcp \
  members add { <cp-node-1>:6443 <cp-node-2>:6443 <cp-node-3>:6443 }

tmsh create ltm virtual k8s-api-vs \
  destination <k8s-api-vip>:6443 ip-protocol tcp \
  pool k8s-api-pool profiles add { tcp }

Repeat the same pair for port 9345 if you want workers to register through the load balancer. Keep the pool members to the control plane nodes only.

Finally, save the configuration with tmsh save sys config. It is the step everyone forgets, and an unsaved F5 loses the virtual servers at the next restart.

Step 9: Validate from the bottom up

Test the platform one layer at a time, starting at the hardware, so a failure points at exactly one layer. Testing only the final URL tells you something is broken and nothing else.

LayerCheckHealthy result
Nodeskubectl get nodes9 nodes Ready
GPUsNode capacity for nvidia.com/gpu8 on each of 6 workers, 48 in total
Drivernvidia-smi in a test pod8 H200 cards listed
Model server/v1/models from inside the clusterThe served model name
GatewaySame call to the LiteLLM ServiceThe public aliases
Entry pointSame call through the F5 nameThe public aliases

GPU visibility

kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.capacity.nvidia\\.com/gpu

A throwaway pod confirms the driver and the runtime work together:

apiVersion: v1
kind: Pod
metadata:
  name: gpu-smoke-test
spec:
  restartPolicy: Never
  tolerations:
    - {key: nvidia.com/gpu, operator: Exists, effect: NoSchedule}
  containers:
    - name: smi
      image: nvcr.io/nvidia/cuda:<tag>-base-ubuntu22.04
      command: ["nvidia-smi"]
      resources:
        limits:
          nvidia.com/gpu: 8

End to end through the F5

This is the request every application will make, so it is the one that matters.

curl https://<llm-fqdn>/v1/chat/completions \
  -H "Authorization: Bearer <virtual-key>" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "standard",
    "messages": [{"role": "user", "content": "Reply with one word: ready"}]
  }'

Then the same call from the official Python client, which is how your applications will use it:

from openai import OpenAI

client = OpenAI(base_url="https://<llm-fqdn>/v1", api_key="<virtual-key>")

stream = client.chat.completions.create(
    model="standard",
    messages=[{"role": "user", "content": "Explain what a leaf switch does."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Watch the output arrive gradually. If the whole answer appears at once after a long pause, something between the model and the client is buffering, and the F5 profile is the first place to look.

Embeddings use the same endpoint:

curl https://<llm-fqdn>/v1/embeddings \
  -H "Authorization: Bearer <virtual-key>" \
  -H "Content-Type: application/json" \
  -d '{"model": "bge-m3", "input": "leaf and spine"}'

Break it on purpose

A platform you have only seen working is not validated. Before real users arrive, run these failure tests.

TestExpected behaviour
Delete one LiteLLM pod during a requestThe other replicas serve traffic and the F5 keeps the pool up
Scale the largest model to zeroRequests fall back to the next tier through the fallbacks rule
Reboot one GPU workerIts models reschedule or fail over, and the other nodes are unaffected
Take one control plane node downkubectl keeps working and the F5 marks that member down
Use a key outside its allowed modelsThe gateway returns an authorisation error, not a model answer

What to watch, and where to take it next

The platform works when every layer has an owner and a health signal. Most of what goes wrong later is a missing signal, not a missing feature.

The failure patterns to expect

SymptomUsual causeWhere to look
GPU node shows 0 GPUsDriver pod not readyLogs of the driver pod in the GPU Operator namespace
Pods stuck in ContainerCreatingCNI not healthycilium status
LoadBalancer Service stays pendingNo address pool or advertisementMetalLB pool and advertisement objects
Model pod restarts on startShared memory too small, or GPU count mismatch/dev/shm volume and parallelism flag
NIM exits at startupNFS not writable by its userExport permissions on the cache path
F5 pool shows down, pods healthyMonitor missing the Host headerMonitor send string
Answers stop mid-sentenceAn idle timeout on one hopF5 TCP profile, then Istio route
Downloads fail with certificate errorsProxy re-signs TLSCorporate CA inside the Job

Where I would take it next

  • Add Redis so per-key rate limits are shared across all gateway replicas.
  • Scrape the DCGM exporter and LiteLLM metrics into the monitoring stack, and alert on GPU memory, queue depth and time to first token, not just pod restarts.
  • Add Kubernetes network policies so only the gateway can reach the model servers. Today anything inside the cluster can call a model directly.
  • Run two replicas of the models that matter most, on different nodes, so a node reboot is not an outage.
  • Rotate the vault credentials and virtual keys on a schedule, and test the rotation before you need it.

The shape of the whole design

The design rests on one idea: keep each layer’s job narrow. The fabric moves bytes, Kubernetes places containers, the operators expose hardware, the model servers run weights, LiteLLM applies policy, and the F5 decides who gets in. When something breaks, that separation is what lets you say which layer, in minutes, instead of arguing about it for a day.

Add a Comment

Your email address will not be published. Required fields are marked *