LOADING
57 words
1 minute
llama.cpp

llama.cpp

Category: Local Runtimes & Chat UIs · GitHub · Docs ⭐ 126,000 · b10717 · MIT · snapshot 2026-08-31 Full profile (deep dive, contenders, matrix): [[AI Tools & Platforms Landscape]]

One-liner: The reference CPU/GPU inference engine and GGUF standard

What problem does it solve?

How does it work?

When would I reach for it?

  • [[Ollama]]
  • [[LocalAI]]
  • [[Exo]]

My exploration notes

Blog post checklist

  • Outline
  • Draft → Blog/posts/tech-explore/llama-cpp/

Some information may be outdated