#local-inference
- Google Releases DiffusionGemma: 26B Open Model with 4× Faster Text Generation Google DeepMind models-llm
- llama.cpp b9754: Real-Time Model Load Progress via SSE and PEG Grammar Parser tools
- llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning llama.cpp (GGML) tools