One API, Four Silicon Families: vLLM on NVIDIA, AMD, TPU, and Intel

Part 2 of “vLLM in 2026: An Infrastructure Engineer’s Field Guide”
The same command, vllm serve <model>, now runs on NVIDIA GPUs, AMD Instinct accelerators, Google TPUs, and Intel hardware (Gaudi, GPUs, and CPUs). For a platform team, that changes procurement, portability, and risk conversations.
But “it runs” is not the same as “it runs equally well.” This post explains how the multi-hardware story works, what actually differs under the hood, and how to think about choosing.
[Add from talks: “NVIDIA and vLLM Full-Stack Collaboration for DeepSeek and MiniMax Performance”, “AMD and vLLM: What’s New”, “Accelerating LLMs: vLLM on Google TPUs”, “Serving vLLM on Intel GPUs, CPUs, and Gaudi”. Speaker benchmarks and supported-feature announcements go in the sections below.]
1. How one engine targets many chips
vLLM separates the engine (scheduler, KV cache manager, API server) from the hardware-specific layer (attention kernels, quantization kernels, collective communication, device memory management). Hardware vendors integrate through a platform abstraction and plugin mechanism, which is why you now see backends maintained by the vendors themselves rather than only by the core team.

The scheduler and PagedAttention logic above the line is shared. Everything below it differs.
2. What differs, by platform
| NVIDIA | AMD | Google TPU | Intel | |
|---|---|---|---|---|
| Software stack | CUDA, cuDNN, NCCL | ROCm, RCCL | XLA, JAX/Pallas kernels | Gaudi software stack, XPU (oneAPI), CPU |
| Scale-up interconnect | NVLink / NVSwitch | Infinity Fabric | ICI (torus topology) | Gaudi on-chip Ethernet (RoCE) |
| Scale-out | InfiniBand / RoCEv2 | Ethernet/RoCE, IB | Data center network (DCN) | Ethernet |
| Low-precision | FP8; FP4 on Blackwell | FP8; newer generations add lower precision | bfloat16, int8/fp8 depending on generation | FP8 on Gaudi; varies elsewhere |
| Strength | Widest kernel and model coverage, earliest access to new features | Large HBM capacity per accelerator | Tight pod-scale interconnect, strong cost story on Google Cloud | Cost/availability alternative; CPU path for small models |
| Watch out for | Cost and supply | Feature parity lags for new model architectures | Different programming and deployment model | Narrower model/feature coverage |
Treat the table as a framework. Generation-specific numbers change fast, so fill in current figures from vendor specs.
NVIDIA: the reference path
New models and optimizations typically land here first. For very large MoE models (DeepSeek-style architectures with Multi-head Latent Attention and many routed experts), the performance story depends on:
- Fused MoE kernels and efficient expert dispatch.
- MLA-aware attention kernels that shrink the KV cache.
- Low-precision execution (FP8, and FP4 on Blackwell-class parts) with the right scaling recipes.
- Communication libraries that make expert-parallel all-to-all efficient (more in Post 3).
[Add from talk: Siyuan Fu’s DeepSeek/MiniMax optimization details and measured speedups.]
AMD: capacity and open software
The headline for AMD Instinct parts is high HBM capacity per device, which directly translates to more KV cache or larger models per node. The ROCm backend has matured to the point where the question is less “does it run?” and more “which features and quantization formats are at parity for my model?”
[Add from talk: Douglas Lehr’s “What’s New”: new hardware support, kernel work, model coverage.]
Google TPU: a different mental model
TPUs are organized in pods connected by a high-bandwidth inter-chip interconnect in a torus topology. Where a GPU cluster gives you NVLink islands stitched together by a fabric, a TPU slice gives you a mesh with different placement rules. The serving path is built around XLA compilation, with kernels written for TPU (e.g., via Pallas). Two practical implications:
- Static shapes matter. XLA compiles for particular shapes, so batching and padding strategy affect compile counts and latency.
- Topology choice is a deployment decision. How you slice the pod determines your parallelism options.
[Add from talk: Qi Zhou’s TPU results: supported models, performance versus GPU, cost framing.]
Intel: breadth of options
Intel’s story spans Gaudi accelerators, Intel GPUs, and CPUs. The CPU path is the interesting one for platform teams: small models, embedding workloads, and dev/test environments can run without any accelerator, which can simplify capacity planning for low-traffic services.
[Add from talk: Chendi Xue’s talk: which model sizes and features are supported per device, and benchmark methodology.]
3. The decision framework
Choose by workload, not by brand:

Then validate with your own numbers. A fair comparison uses the same model, same quantization where possible, same prompt/output length distribution, and reports:
- Tokens/sec at a fixed p99 latency target (not peak throughput).
- Cost per million output tokens, including power and networking.
- Time to first token at realistic concurrency.
- Engineering cost: how much custom work was needed?
4. Portability traps
- Quantization formats are not universal. A checkpoint quantized for one backend’s kernels may not run on another.
- Attention backends differ, so numerics can differ slightly. Re-run your eval suite per platform.
- Collectives differ (NCCL vs RCCL vs ICI). Multi-node tuning knowledge does not transfer one-to-one.
- Container images and drivers become a matrix to maintain: OS, driver, runtime, vLLM version, per hardware family.
Design-review checklist
- Is the multi-vendor requirement real (supply risk, cost leverage) or theoretical?
- Have we benchmarked on our model, quantization, and traffic shape, not vendor demos?
- Do we have an eval suite to catch numerical differences per platform?
- Is the image/driver matrix owned by someone?
- Do we compare on cost per million tokens at a p99 target?
Next: Post 3: Distributed Inference Is a Network Problem
Planning Private or Sovereign AI Infrastructure?
Building an enterprise LLM platform requires more than selecting GPUs.
Your GPU compute, high-performance network fabric, Kubernetes platform, security controls, model-serving architecture, observability and application access layer need to work together.
TREZBON helps organisations design and implement secure, scalable infrastructure for modern enterprise and AI workloads.
Whether you are evaluating private LLM deployment, GPU infrastructure, enterprise networking, cybersecurity or data-centre architecture, our team can help you move from design to implementation.
Discuss your infrastructure requirements with TREZBON:
www.trezbon.com/#contact