Skip to content
AI Primer
TOOL6 stories

llama.cpp

Stories, products, and related signals connected to this tag in Explore.

WORKFLOW4w ago
Qwen3.8-Flash-Next runs from SSD on M4 Max at 40 tokens per second

A developer reports streaming 60% of Qwen3.8-Flash-Next experts from disk on demand, running the full model in 37 GB at 40 tokens per second. BF16 and GGUF weights are also available for local deployments.

WORKFLOW1mo ago
Users report Qwen 3.8 27B agents vary sharply by harness

Local users report that Qwen 3.8 27B agent behavior changes substantially with the harness, quantization, context settings, and hardware. One 3-bit MacBook Air run at 57K context took 63 hours.

WORKFLOW1mo ago
Qwen 122B runs on older laptop with llama.cpp in user test

A LocalLLaMA post ran a 122B Qwen model on an older laptop with llama.cpp. The run had very long load and generation times, while another report put Qwen 3.6 35B at 21 tok/s on a Radeon 7600 after ROCm tuning.

WORKFLOW2mo ago
LocalLLaMA users report near-6x Qwen 3.6 27B speedups with MTP

A LocalLLaMA benchmark on Qwen 3.6 27B and RTX 6000 PRO reports near-6x speedups from MTP, DFlash, and n-gram drafting. Related tests cover remote prefill, NUMA offload, GPU clock tuning, and llama.cpp Gemma 4 support.

RELEASE2mo ago
Tencent releases 1-bit and 4-bit GGUF weights for 295B Hy3 single-GPU runs

Tencent released 1-bit and 4-bit GGUF builds for its 295B Hy3 model with llama.cpp support and MTP. Posts cite 88–92GB local runs and SWE-Bench scores of 75.4% Verified and 53.9% Pro.

RELEASE4mo ago
llama.cpp launches official site with one-line installer and unified `llama` CLI

llama.cpp now has an official website and a single-line installer that provides one `llama` entrypoint for running, serving, and agent integrations. The packaging change simplifies local setup while reusing GGUF models already on disk.

AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.