vLLM v0.27.0 Ships Kimi K3 Support and DeepSeek-V4 Optimizations
vLLM Project
vLLM 0.27.0 (561 commits, 242 contributors) adds full-stack Kimi K3 support, new Qwen3.5 and K-EXAONE-2.0-750B-A37B model support, DeepSeek-V4 routing-kernel optimizations, a PyTorch 2.13.0 upgrade, FlashAttention 4 on SM100 with FP8 KV cache, Model Runner V2 for embedding/classification workloads, and early NVIDIA Rubin and ROCm gfx1250 hardware support.
Why it matters
A major release expanding vLLM's model coverage and inference performance for large-scale self-hosted serving.
Importance: 3/5
Notable release: major version bump with broad new model and hardware support, single official source.
Sources
official
vLLM releases