tool listing
llama.cpp
The C and C++ inference engine that most local model runners are built on. Runs quantised models on ordinary CPUs and consumer GPUs.
- Site
- github.com/ggml-org/llama.cpp
- Category
- local
- Pricing
- free, open source (MIT)
- Last verified
- Wiki entry
- llama.cpp