tool listing

llama.cpp

The C and C++ inference engine that most local model runners are built on. Runs quantised models on ordinary CPUs and consumer GPUs.

Site
github.com/ggml-org/llama.cpp
Category
local
Pricing
free, open source (MIT)
Last verified
Wiki entry
llama.cpp