#serving
- MinT: Managed Infrastructure for Training and Serving Millions of LLMs Mind Lab research
- vLLM v0.25.0: Model Runner V2 Default, PagedAttention Retired, Transformers Backend Parity tools
- vLLM Adds Day-0 Support for MiniMax M3 Open Weights with 1M-Context Sparse Attention MiniMax tools
- vLLM v0.24.0: Model Runner V2 Default, Rust Frontend, SM90 FP8 Speedups vLLM tools
- vLLM v0.27.0 Ships Kimi K3 Support and DeepSeek-V4 Optimizations vLLM Project tools
- Modal Launches Auto Endpoints for Production-Grade Open-Model LLM Inference Modal tools
- ELDR: Expert-Locality-Aware Routing Cuts MoE Serving Latency by up to 14% Microsoft Research research
- Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler Dharma-AI tools
- Groq closes $350M Series A to build an AI inference cloud Groq industry
- SGLang v0.5.18 lands major model and perf updates tools
- SGLang v0.5.18 SGLang tools