llama.cpp

Other Free

Efficient inference engine for running LLMs locally.

About llama.cpp

llama.cpp enables CPU/GPU local model execution with quantization support.