tool · tool/llama-cpp

llama.cpp

Facts

license
MITsource, accessed 2026-08-28
stated goal
LLM and VLM inference in C/C++ with minimal setup across a wide range of hardwaresource, accessed 2026-08-28
backends
more than fifteen backends, including CUDA, HIP, Metal, Vulkan, SYCL and OpenVINOsource, accessed 2026-08-28
quantization types
k-quants Q2_K through Q6_K plus the IQ series, each with a measured bits-per-weight figure in the repository's own tablesource, accessed 2026-08-28

Timeline

  1. k-quants merged, introducing super-block quantization types Q2_K through Q6_Ksource