#cuda
- Qualcomm Acquires Modular for $3.92B to Challenge CUDA Lock-in Qualcomm industry
- Hugging Face Transformers: Async Continuous Batching Achieves 22% Inference Speedup Hugging Face tools
- llama.cpp v0.3.0 — first tagged stable release of the b10621 nightly; dots3-note multimodal, GLM-4.5-Air MTP, DeepSeek 4 tensor-split, ggml v0.22.0 ggml-org tools
- llama.cpp b9589–b9592: CUDA SSM Sync Fix and Mamba Memory Optimization tools
- Ollama v0.31.2: MLX small-batch matmul kernel, llama.cpp build 9840, CUDA updates tools
- Dharma-AI lifts GPU-cluster utilization 53.6% → 87.0% by encoding physical constraints into the scheduler Dharma-AI tools