Data Science Wire

vLLM Reaches 25K Total TPS/GPU on Qwen3.5

vLLM BlogAug 64 min read

How vLLM reaches 25K total TPS/GPU on Qwen3.5-397B-A17B-NVFP4 with GB200 NVL72 disaggregated serving, Blackwell GDN kernels, HMA cache transfer, async scheduling fixes, and srt-slurm recipes.

Read the full story at vLLM Blog

More in MLOps / LLMOps