llama.cpp
LLM inference in C/C++
Open-source C/C++ runtime and library for running LLM inference locally or in the cloud, including CLI and OpenAI-compatible server usage, GGUF model workflows, quantization, and many CPU/GPU hardware backends.

Recent stories
0 linked stories
No linked stories yet.