llama.cpp ships v0.1.2 pre-release plus daily b10483–b10488 with SYCL perf fix, OpenVINO bump, DGX Spark CUDA tuning

llama.cpp (GGML)

Tools official 1 src. ~1 min

Aug 18 work spans v0.1.2 pre-release (ggml sync to 0.20.2, MCP stdio docs + CORS defaults, integer tokenizer scores, refactor of Built-In Tools naming), b10488 (OpenVINO 2026.3), b10486 (LFM2 image tiling threshold fix), b10485 (ggml sync), and b10483 (cmake vendor:: alias targets). Aug 17 added b10472 (CUDA skip UMA override for HIP, #18159), b10470 (release.yml pushes tag explicitly), b10456 (SYCL q4_0→f32 20.21→158.19 GB/s on Arc 70), and b10455 (SYCL OPT_STEP_ADAMW/SGD).

Why it matters

The SYCL quantized-copy kernel fix delivers an ~8x throughput improvement for q4_0→f32 on Arc GPUs — a single kernel change that materially changes what local inference looks like on Intel discrete cards. OpenVINO 2026.3 and DGX Spark CUDA MMVQ tuning also widen the deployment matrix.

Importance: 2/5

default

Sources