Skip to content
AI Primer

llama.cpp

LLM inference in C/C++

Open-source C/C++ runtime and library for running LLM inference locally or in the cloud, including CLI and OpenAI-compatible server usage, GGUF model workflows, quantization, and many CPU/GPU hardware backends.

Screenshot of llama.cpp website

Recent stories

0 linked stories
No linked stories yet.
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.