vLLM Reaches 25K Total TPS/GPU on Qwen3.5
vLLM BlogAug 64 min read
How vLLM reaches 25K total TPS/GPU on Qwen3.5-397B-A17B-NVFP4 with GB200 NVL72 disaggregated serving, Blackwell GDN kernels, HMA cache transfer, async scheduling fixes, and srt-slurm recipes.
