One API, Four Silicon Families: vLLM on NVIDIA, AMD, TPU, and Intel

Part 2 of “vLLM in 2026: An Infrastructure Engineer’s Field Guide”


The same command, vllm serve <model>, now runs on NVIDIA GPUs, AMD Instinct accelerators, Google TPUs, and Intel hardware (Gaudi, GPUs, and CPUs). For a platform team, that changes procurement, portability, and risk conversations.

But “it runs” is not the same as “it runs equally well.” This post explains how the multi-hardware story works, what actually differs under the hood, and how to think about choosing.

[Add from talks: “NVIDIA and vLLM Full-Stack Collaboration for DeepSeek and MiniMax Performance”, “AMD and vLLM: What’s New”, “Accelerating LLMs: vLLM on Google TPUs”, “Serving vLLM on Intel GPUs, CPUs, and Gaudi”. Speaker benchmarks and supported-feature announcements go in the sections below.]


1. How one engine targets many chips

vLLM separates the engine (scheduler, KV cache manager, API server) from the hardware-specific layer (attention kernels, quantization kernels, collective communication, device memory management). Hardware vendors integrate through a platform abstraction and plugin mechanism, which is why you now see backends maintained by the vendors themselves rather than only by the core team.

The scheduler and PagedAttention logic above the line is shared. Everything below it differs.

2. What differs, by platform

NVIDIAAMDGoogle TPUIntel
Software stackCUDA, cuDNN, NCCLROCm, RCCLXLA, JAX/Pallas kernelsGaudi software stack, XPU (oneAPI), CPU
Scale-up interconnectNVLink / NVSwitchInfinity FabricICI (torus topology)Gaudi on-chip Ethernet (RoCE)
Scale-outInfiniBand / RoCEv2Ethernet/RoCE, IBData center network (DCN)Ethernet
Low-precisionFP8; FP4 on BlackwellFP8; newer generations add lower precisionbfloat16, int8/fp8 depending on generationFP8 on Gaudi; varies elsewhere
StrengthWidest kernel and model coverage, earliest access to new featuresLarge HBM capacity per acceleratorTight pod-scale interconnect, strong cost story on Google CloudCost/availability alternative; CPU path for small models
Watch out forCost and supplyFeature parity lags for new model architecturesDifferent programming and deployment modelNarrower model/feature coverage

Treat the table as a framework. Generation-specific numbers change fast, so fill in current figures from vendor specs.

NVIDIA: the reference path

New models and optimizations typically land here first. For very large MoE models (DeepSeek-style architectures with Multi-head Latent Attention and many routed experts), the performance story depends on:

  • Fused MoE kernels and efficient expert dispatch.
  • MLA-aware attention kernels that shrink the KV cache.
  • Low-precision execution (FP8, and FP4 on Blackwell-class parts) with the right scaling recipes.
  • Communication libraries that make expert-parallel all-to-all efficient (more in Post 3).

[Add from talk: Siyuan Fu’s DeepSeek/MiniMax optimization details and measured speedups.]

AMD: capacity and open software

The headline for AMD Instinct parts is high HBM capacity per device, which directly translates to more KV cache or larger models per node. The ROCm backend has matured to the point where the question is less “does it run?” and more “which features and quantization formats are at parity for my model?”

[Add from talk: Douglas Lehr’s “What’s New”: new hardware support, kernel work, model coverage.]

Google TPU: a different mental model

TPUs are organized in pods connected by a high-bandwidth inter-chip interconnect in a torus topology. Where a GPU cluster gives you NVLink islands stitched together by a fabric, a TPU slice gives you a mesh with different placement rules. The serving path is built around XLA compilation, with kernels written for TPU (e.g., via Pallas). Two practical implications:

  • Static shapes matter. XLA compiles for particular shapes, so batching and padding strategy affect compile counts and latency.
  • Topology choice is a deployment decision. How you slice the pod determines your parallelism options.

[Add from talk: Qi Zhou’s TPU results: supported models, performance versus GPU, cost framing.]

Intel: breadth of options

Intel’s story spans Gaudi accelerators, Intel GPUs, and CPUs. The CPU path is the interesting one for platform teams: small models, embedding workloads, and dev/test environments can run without any accelerator, which can simplify capacity planning for low-traffic services.

[Add from talk: Chendi Xue’s talk: which model sizes and features are supported per device, and benchmark methodology.]

3. The decision framework

Choose by workload, not by brand:

Then validate with your own numbers. A fair comparison uses the same model, same quantization where possible, same prompt/output length distribution, and reports:

  • Tokens/sec at a fixed p99 latency target (not peak throughput).
  • Cost per million output tokens, including power and networking.
  • Time to first token at realistic concurrency.
  • Engineering cost: how much custom work was needed?

4. Portability traps

  • Quantization formats are not universal. A checkpoint quantized for one backend’s kernels may not run on another.
  • Attention backends differ, so numerics can differ slightly. Re-run your eval suite per platform.
  • Collectives differ (NCCL vs RCCL vs ICI). Multi-node tuning knowledge does not transfer one-to-one.
  • Container images and drivers become a matrix to maintain: OS, driver, runtime, vLLM version, per hardware family.

Design-review checklist

  • Is the multi-vendor requirement real (supply risk, cost leverage) or theoretical?
  • Have we benchmarked on our model, quantization, and traffic shape, not vendor demos?
  • Do we have an eval suite to catch numerical differences per platform?
  • Is the image/driver matrix owned by someone?
  • Do we compare on cost per million tokens at a p99 target?

Next: Post 3: Distributed Inference Is a Network Problem

Planning Private or Sovereign AI Infrastructure?

Building an enterprise LLM platform requires more than selecting GPUs.

Your GPU compute, high-performance network fabric, Kubernetes platform, security controls, model-serving architecture, observability and application access layer need to work together.

TREZBON helps organisations design and implement secure, scalable infrastructure for modern enterprise and AI workloads.

Whether you are evaluating private LLM deployment, GPU infrastructure, enterprise networking, cybersecurity or data-centre architecture, our team can help you move from design to implementation.

Discuss your infrastructure requirements with TREZBON:
www.trezbon.com/#contact

Add a Comment

Your email address will not be published. Required fields are marked *