llama.cpp
An open-source C/C++ engine for running language models locally on CPUs and GPUs, and the basis of many local AI apps.
About llama.cpp
Running models on laptops and small servers needs an efficient, dependency-light runtime.
Who it’s for
- Developers running models locally
- Builders of local AI apps
When to consider it
Consider llama.cpp when you want to run quantized models on your own hardware with minimal setup.
Tradeoffs & limitations
- MIT licensed; it is a library and command-line tool rather than a consumer app.
SOTA overview · Documentation-based assessment · Sources & review method
Updates
No updates shared yet.
Discussion
Newest firstAsk a question or share how you use llama.cpp.
Keep it helpful. Community rules
Loading discussion…
Comparisons & guides
Sources & review method
Documentation-based assessment · Oct 2, 2026 · Prepared with AI assistance; not a hands-on benchmark.
Checked by SOTA · AI-assisted documentation review. Selection advice is our assessment; verify current requirements for your deployment.
Import history & original evidence
llama.cpp on GitHub
An open-source C/C++ engine for running language models locally on CPUs and GPUs, and the basis of many local AI apps.
Based on official pages and announcements checked on 2026-10-02. No hands-on test, performance benchmark or popularity ranking is claimed.