llama.cpp

llama.cpp

llama.cpp is a high-performance LLM inference engine written in C/C++ with minimal dependencies.

FreeWebiOSAndroidOpen sourceKoreanLLM-basedMultimodal
Visit websitegithub.com
Compare with book-to-skillExplore llama.cpp alternatives

Overview

llama.cpp is a high-performance LLM inference engine written in C/C++ with minimal dependencies. It supports various hardware accelerations including Apple Silicon's Metal and NVIDIA CUDA, enabling large models to run efficiently on consumer-grade hardware using the GGUF format.

Key features

  • Pure C/C++ implementation
  • GGUF format support
  • Advanced quantization
  • Apple Silicon Metal optimization
  • CUDA/OpenCL acceleration
  • Built-in HTTP server
  • Multiple language bindings

Pricing

FreeStarting price: Free
View pricing page

Verified on:

Use cases

  • Running local LLMs on personal PCs
  • Embedding AI models in mobile devices
  • Building offline AI assistants

Who it is for

Individual DevelopersLocal AI ResearchersEmbedded Systems Engineers

Integrations

PythonNode.jsRustGo

Tags

GGUF

How we verified this

Company, pricing, and feature details come from the primary sources below and our latest verification pass. When sources disagree, the official source and the most recent check win.

Last verified 08/30/2026Verified sources: 1

Alternatives

Tools you can use instead