-
JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
Hao AI Lab, UC San Diego
research
-
DeepSeek Open-Sources DSpark: 57–85% Inference Speedup for V4 in Production
DeepSeek
tools
-
JetSpec: Parallel Tree Drafting Achieves 9.64× Speculative Decoding Speedup
Hao AI Lab, UCSD
research
-
SGLang v0.5.11: Speculative Decoding V2 as Default and Eight New Model Architectures
tools
-
Orthrus: 7.8x Inference Speedup for Qwen3 via Autoregressive-Diffusion KV Sharing
research
-
vLLM v0.21.0: Blackwell MLA Backend, HMA KV Offload, Spec Decode for Reasoning Models
vLLM Project
tools
-
BlockPilot: Instance-Adaptive Block Size for Diffusion-Based Speculative Decoding
research
-
Ollama v0.23.1: Gemma 4 MTP Speculative Decoding Delivers 2× Speed on Apple Silicon
tools
-
llama.cpp June 16 Builds: Eagle3 Speculative Decoding, Vulkan UMA Memory, NVFP4 Fixes
tools
-
llama.cpp Builds b9830–b9837: DFlash v2, MiniCPM5 Parser, --reasoning-preserve Flag
ggml-org
tools
-
llama.cpp Adds Tencent Hunyuan 3, Minimax2 Eagle3 Speculative Decoding, and SYCL Fused MoE
tools