llama.cpp v0.6.0: Metal Speedups, New Models and APIs
llama.cpp v0.6.0 was released on October 5, 2026 with a new extended batch API, fresh model support, speculative decoding improvements and major Apple Metal work. The release matters most to developers running local AI, maintaining inference servers or optimizing Apple Silicon workloads.
llama.cpp v0.6.0 highlights
- New
llama_batch_extAPI withllama_process(). - Support for GLM-5.3-Flash and the Clef decision model.
- MTP speculative decoding for Qwen4Exp.
- New
/v1/systemoneserver API for decision models. - New Metal flash-attention and few-row MMA kernels.
- ggml updated to v0.26.0.
The official llama.cpp v0.6.0 release notes are the primary source for these changes.
A new extended batch API
The new llama_batch_ext interface supports mixed token and embedding inputs and per-token state embeddings used by newer multi-token-prediction and deep-stack architectures. That is an architectural change for native integrations even if applications using higher-level server endpoints do not need to rewrite immediately.
GLM-5.3-Flash and Clef support
The release adds GLM-5.3-Flash, described by the project as a 320B text-and-vision hybrid model. It also adds full text and vision support for the Clef decision model. Model support in llama.cpp matters because many local-AI applications depend on the runtime rather than implementing each architecture themselves.
Qwen4Exp gets MTP speculative decoding
The release notes report about a 1.5x decode speedup on DGX Spark for the tested Qwen4Exp MTP speculative-decoding path. That result is workload-specific and should not be treated as a guarantee for every GPU, model or context size.
Apple Metal improvements
Version 0.6.0 adds a tensor-API flash-attention kernel for F16 KV data and new few-row MMA matrix-multiplication kernels for speculative and batched decoding. The project reports up to about 3x faster matmul for the targeted Apple GPU paths. That does not mean every end-to-end generation becomes three times faster; actual gains depend on the model, batch shape, quantization and hardware.
New server API for decision models
The llama.cpp server gains a /v1/systemone endpoint for several decision-model families. This gives applications a more purpose-built interface for local scoring or action selection instead of forcing every task through a chat-completions pattern.
What should local Mac users benchmark?
- Prompt-processing speed at the same context length.
- Single-user decode speed.
- Batched decoding with multiple sequences.
- Speculative decoding where supported.
- Peak memory and KV-cache behavior.
- The exact quantization used in production.
For a useful before/after comparison, keep the model file, context size and runtime flags constant.
How to upgrade safely
- Record your current build and benchmark results.
- Read the v0.6.0 release notes for API changes touching your integration.
- Build with the same backend options you already use.
- Run known models and prompts before enabling new features.
- Compare output correctness, memory use and throughput.
Because llama.cpp changes quickly, production teams should pin a reproducible version instead of deploying an arbitrary development commit.
FAQ
When was llama.cpp v0.6.0 released?
The tagged release was published on October 5, 2026.
Does it support GLM-5.3-Flash?
Yes. The release notes list GLM-5.3-Flash as a new supported text-and-vision hybrid model.
Is llama.cpp now three times faster?
Not universally. The up-to-3x figure applies to targeted few-row Metal matmul paths used in speculative and batched decoding.
What is the new batch API?
llama_batch_ext with llama_process() supports mixed token and embedding batches plus per-token state embeddings.
