llama.cpp
by Various
LLM inference in C/C++
Apps
llama.cpp
Added 1 June 2026
Overview
llama.cpp runs large language models locally using C/C++ inference optimized for CPU and GPU execution. It enables developers to deploy quantized models with minimal dependencies and memory overhead, making LLM inference practical on consumer hardware.
Best for
Best for
Developers building privacy-first applications or deploying models on resource-constrained devices
Use cases
- Running open-source models offline without API calls
- Embedding LLM capabilities into applications with low latency
- Quantizing and optimizing models for edge deployment
Notes
llama.cpp runs large language models locally using C/C++ inference optimized for CPU and GPU execution. It enables developers to deploy quantized models with minimal dependencies and memory overhead, making LLM inference practical on consumer hardware.
114,160 stars on GitHub. Last updated 2026-06-01. Licensed MIT.
Use cases
- Running open-source models offline without API calls
- Embedding LLM capabilities into applications with low latency
- Quantizing and optimizing models for edge deployment
Pros
- Extremely efficient inference on CPU and GPU with minimal resource requirements
- Supports quantized model formats, reducing model size by 4-8x without major quality loss
- Active community with broad hardware compatibility and regular model support updates
Cons
- Steeper setup curve than API-based solutions, requires compilation and model management
- Performance varies significantly based on hardware, CPU inference is substantially slower than GPU
- Limited to inference only, no built-in fine-tuning or training capabilities
Indexed from awesome-generative-ai and enriched against its public facts.
Pros
- Extremely efficient inference on CPU and GPU with minimal resource requirements
- Supports quantized model formats, reducing model size by 4-8x without major quality loss
- Active community with broad hardware compatibility and regular model support updates
Cons
- Steeper setup curve than API-based solutions, requires compilation and model management
- Performance varies significantly based on hardware, CPU inference is substantially slower than GPU
- Limited to inference only, no built-in fine-tuning or training capabilities
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
vLLM
Community
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang
Community
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT-LLM
Community
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NV
LMDeploy
Community
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
MNN-LLM
Community
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
exllama
Community
A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
Llama 3-8|70B
Community
[Llama 2-7 13 70B](https://llama.meta.com/llama2/)
LMDeploy
Community
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
mistral.rs
Community
Fast, flexible LLM inference
MNN-LLM
Community
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
Rapid-MLX
Community
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Dr
Shimmy
Community
⚡ Python-free Rust inference server — OpenAI-API compatible. GGUF + SafeTensors, hot model swap, auto-discovery, single binary. FREE now, FREE forever.
text-generation-inference
Community
Large Language Model Text Generation Inference
bitnet.cpp
Various
Official inference framework for 1-bit LLMs
gpt4all
Various
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
OpenAI API
Various
Announcement of the OpenAI API for text-to-text general-purpose AI models based on GPT-3. OpenAI blog, June 11, 2020.
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
ollama
Community
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Open WebUI
Various
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
gpt4all
Various
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
Private GPT
Community
Interact with your documents using the power of GPT, 100% privately, no data leaks
Anything LLM
Community
The all-in-one AI productivity accelerator. On device and privacy first with no annoying setup or configuration.
Clippy
Community
AI programming assistant
openinterpreter
Community
Interpreter lets you work alongside agents that can edit your documents, fill PDF forms, and more.
wispy
Community
Chrome extension
inference.sh
ac.inference.sh
Run 150+ AI apps — image, video, audio, LLMs, 3D and more. Browse, execute, stream results.
mediar-ai/screenpipe
Various
YC (S26) | AI that knows what you've seen, said, or heard. Records everything you do, say, hear 24/7, local, private, secure
BenyD/haypile
Various
Private search and Q&A for your documents. One binary that watches your folders and answers questions with citations, fully local.
Agent-LLM
Community
AGiXT is a dynamic AI Agent Automation Platform that seamlessly orchestrates instruction management and complex task execution across diverse AI providers. Combining adaptive memor
Anything LLM
Community
The all-in-one AI productivity accelerator. On device and privacy first with no annoying setup or configuration.
AutoGPT
Community
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
BlockAGI
Community
Your Self-Hosted, Hackable Research Agent Inspired by AutoGPT
Casibase
Community
⚡️next-generation personal AI assistant powered by LLM, RAG and agent loops, supporting computer-use, browser-use and coding agent, demo: https://demo.openagentai.org
CodeQwen1.5-7B
Community
GITHUB HUGGING FACE MODELSCOPE DEMO DISCORD Introduction The advent of advanced programming tools, which harnesses the power of large language models (LLMs), has significantly en
deploy-llms-with-ansible
Community
Easily deploy LLMs with Ansible. Uses Docker with llama.cpp or ollama. Secured with whitelisted IPs.
Flappy
Community
Production-Ready LLM Agent SDK for Every Developer
Knowledge GPT
Community
Accurate answers and instant citations for your documents.
LangChain
Community
The agent engineering platform.
Llama 3-8|70B
Community
[Llama 2-7 13 70B](https://llama.meta.com/llama2/)
LLama Cpp Agent
Community
The llama-cpp-agent framework is a tool designed for easy interaction with Large Language Models (LLMs). Allowing users to chat with LLM models, execute structured function calls a
LLMKube
Community
Kubernetes operator for local LLM inference with llama.cpp, vLLM, TGI, and mlx-server — multi-GPU NVIDIA + Apple Silicon Metal, autoscaling, air-gapped, production-ready
LLocalSearch
Community
LLocalSearch is a completely locally running search aggregator using LLM Agents. The user can ask a question and the system will use a chain of LLMs to find the answer. The user ca
LMQL
Community
Language Model Query Language
Local GPT
Community
Chat with your documents on your local device using GPT models. No data leaves your device and 100% private.
Off Grid
Community
The Swiss Army Knife of Offline AI. Chat, Speak, and Generate Images - Privacy First, Zero Internet. Download an LLM and use it on your mobile device. No data ever leaves your phon
ollama
Community
Get up and running with Kimi-K2.5, GLM-5, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Private GPT
Community
Interact with your documents using the power of GPT, 100% privately, no data leaks
QA-Pilot
Community
QA-Pilot is an interactive chat project that leverages online/local LLM for rapid understanding and navigation of GitHub code repository.
Rigging
Community
Lightweight LLM Interaction Framework
Scale Spellbook
Community
Accelerate and scale Generative AI across your enterprise with the platform to transform your data into customized enterprise-ready Generative AI applications.
Serge
Community
A web interface for chatting with Alpaca through llama.cpp. Fully dockerized, with an easy to use API.
SuperAGI
Community
SuperAGI - A dev-first open source autonomous AI agent framework. Enabling developers to build, manage & run useful autonomous agents quickly and reliably.
tabby
Community
Self-hosted AI coding assistant
AutoGen
Various
A programming framework for agentic AI
Chatbot UI
Various
Chatbot UI
gpt4all
Various
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
Harbor
Various
Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore.
Jan
Various
Jan is an open-source alternative to ChatGPT. Run open-source AI models locally or connect to cloud models like GPT, Claude and others.
Kilo
Various
Kilo is the open source AI coding agent for VS Code, JetBrains, CLI, and Cloud. Access 500+ models, bring your own keys at zero markup, and keep code private with local models.
LangChain
Various
LangChain provides the engineering platform and open source frameworks developers use to build, test, and deploy reliable AI agents.
LM Studio
Various
Run local AI models like gpt-oss, Llama, Gemma, Qwen, and DeepSeek privately on your computer.
LLM
Various
LLM: A CLI utility and Python library for interacting with Large Language Models
Local Deep Research
Various
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Every
Msty
Various
Msty AI builds local-first AI tools: Msty Studio is your private AI workspace, and Msty Claw runs autonomous multi-step tasks with sandboxed control on your machine.
Open Interpreter
Various
A natural language interface for computers
PyGPT
Various
PyGPT is an open‑source desktop AI assistant for Windows, macOS and Linux. Chat, agents, web search, run Python, TTS/STT, plugins, long‑term memory.
quivr
Various
Opiniated RAG for integrating GenAI in your apps 🧠 Focus on your product rather than the RAG. Easy integration in existing products with customisation! Any LLM: GPT4, Groq, Llama.
Stable Horde
Various
AI Horde
Teleprompter
Various
An on-device AI for your meetings that listens to you and makes charismatic quote suggestions.
TurboPilot
Various
Turbopilot is an open source large-language-model based code completion engine that runs locally on CPU
Vibe Transcribe
Various
Local-first transcription for audio and video with AI summaries, multilingual support, and privacy-focused processing.
prima.cpp
Community
A distributed implementation of llama.cpp that lets you run 70B-level LLMs on your everyday devices.
Wllama
Community
WebAssembly binding for llama.cpp - Enabling on-browser LLM inference
LM Studio
Various
Run local AI models like gpt-oss, Llama, Gemma, Qwen, and DeepSeek privately on your computer.
privateGPT
Various
Interact with your documents using the power of GPT, 100% privately, no data leaks
cameronrye/openzim-mcp
Various
OpenZIM MCP is a modern, secure, and high-performance MCP (Model Context Protocol) server that enables AI models to access and search ZIM format knowledge bases offline.
Jwrede/llmprobe
Various
Synthetic monitoring and CI smoke tests for LLM inference endpoints.
ocbenji/bitcoinbenji-mcp
Various
MCP server for the Bitcoin Benji API — Lightning-paid Bitcoin mempool intelligence + sovereign on-prem AI inference (L402). No third-party APIs.
Agency
Community
🕵️♂️ Library designed for developers eager to explore the potential of Large Language Models (LLMs) and other generative AI through a clean, effective, and Go-idiomatic approach.
AgentField
Community
Build, run and scale AI agents like API and microservices - observable,auditable and identity-aware from day one.
AI Gateway
Community
A blazing fast AI Gateway with integrated guardrails. Route to 1,600+ LLMs, 50+ AI Guardrails with 1 fast & friendly API.
Awesome-Code-LLM
Community
👨💻 An awesome and curated list of best code-LLM for research.
awesome-japanese-llm
Community
日本語LLMまとめ - Overview of Japanese LLMs
Awesome-LLM-Inference
Community
📖A curated list of Awesome LLM/VLM Inference Papers with codes: WINT8/4, FlashAttention, PagedAttention, MLA, Parallelism, etc. 🎉🎉
Baichuan-7|13B
Community
AGI Large Language Models
BLOOMZ&mT0
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Chainlit
Community
A Python library for making chatbot interfaces.
Codestral-7|22B
Community
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
CompassRank
Community
评测榜单旨在为大语言模型和多模态模型提供全面、客观且中立的得分与排名,同时提供多能力维度的评分参考,以便用户能够更全面地了解大模型的能力水平。
Cursor
Community
Built to make you extraordinarily productive, Cursor is the best coding agent.
DeepSeek-R1
Community
First-generation reasoning models from DeepSeek.
DeepSeek-Math-7B
Community
DeepSeek Math series
DeepSeek-V2.5
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
DeepSeek-VL-1.3|7B
Community
DeepSeek-VL model series
Gemma
Community
Checking your browser - reCAPTCHA
Gemma2-9|27B
Community
Gemma 2, our next generation of open models, is now available globally for researchers and developers.
GLM-2|6|10|13|70B
Community
Org profile for THUDM on Hugging Face, the AI community building the future.
Google "We Have No Moat, And Neither Does OpenAI"
Community
Leaked Internal Google Document Claims Open Source AI Will Outcompete Google and OpenAI
Grok-1-314B-MoE
Community
Grok-1-314B-MoE — indexed from awesome-llm
Guidance
Community
A guidance language for controlling large language models.
InternLM2-1.8|7|20B
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
LazyLLM
Community
Easiest and laziest way for building multi-agent LLMs applications.
lighteval
Community
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Llama 3.2-1|3|11|90B
Community
[Llama 3.1-8 70 405B](https://llama.meta.com/)
Llama 3-8|70B
Community
[Llama 2-7 13 70B](https://llama.meta.com/llama2/)
LLaMA Cult and More
Community
Large Language Models for All, 🦙 Cult and More, Stay in touch !
llm-ui
Community
The React library for LLMs
LMQL
Community
Language Model Query Language
MemFree
Community
MemFree - Hybrid AI Search Engine & AI Page Generator
MemGPT
Community
Letta is the platform for building stateful agents: AI with advanced memory that can learn and self-improve over time.
MiniChain
Community
A tiny library for coding with large language models.
MiniCPM-2B
Community
The MiniCPM family of LLMs and VLLMs.
Mixtral-8x7B
Community
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
ModelFusion
Community
The TypeScript library for building AI applications.
Moonlight-A3B
Community
Moonshot's Compute-efficient MoE LLM, first Scaling Up of Muon Optimizer
OLMo-7B
Community
Artifacts for the first set of OLMo models.
Outlines
Community
Structured Outputs
pgvector
Community
Open-source vector similarity search for Postgres
Phi1-1.3B
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Phi3-3.8|7|14B
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Qwen-1.8B|7B|14B|72B
Community
Qwen - a Qwen Collection
Qwen2-0.5B|1.5B|7B|57B-A14B-MoE|72B
Community
GITHUB HUGGING FACE MODELSCOPE DEMO DISCORD Introduction After months of efforts, we are pleased to announce the evolution from Qwen1.5 to Qwen2. This time, we bring to you: Pret
Qwen2-Math-1.5B|7B|72B
Community
GITHUB HUGGING FACE MODELSCOPE DISCORD 🚨 This model mainly supports English. We will release bilingual (English and Chinese) math models soon. Introduction Over the past year, w
RecurrentGemma-2B
Community
Open weights language model from Google DeepMind, based on Griffin.
Robocorp
Community
Create 🐍 Python AI Actions and 🤖 Automations, and deploy & operate them anywhere
RWKV-v4|5|6
Community
Org profile for RWKV on Hugging Face, the AI community building the future.
Semantic Kernel
Microsoft
Microsoft's enterprise-flavoured framework for AI agents. .NET-first, with Python and Java siblings.
simple-evals
Community
Eval tools by OpenAI.
StableLM-3B
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
StarCoder-1|3|7B
Community
All models, datasets, and demos related to StarCoder!
Swiss Army Llama
Community
A FastAPI service for semantic text search using precomputed embeddings and advanced similarity measures, with built-in support for various file types through textract.
unslothai
Community
Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.
Yi-34B
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
DeepSeek
Various
Org profile for DeepSeek on Hugging Face, the AI community building the future.
GitHub Models
Various
Find and experiment with AI models to develop a generative AI application.
Grok
Various
An LLM by xAI with [open source](https://github.com/xai-org/grok-1) and open weights. #opensource
LibreChat
Various
LibreChat brings together all your AI conversations in one unified, customizable interface.
LLaMA
Various
Llama LLM, a foundational, 65-billion-parameter large language model by Meta. Meta, February 23rd, 2023. #opensource
LLM Stats
Various
The LLM Leaderboard — independent ranking of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models by intelligence, speed and price. Composite LLM Stats Score updated continuou
Mistral
Various
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Open LLMs
Various
📋 A list of open LLMs available for commercial use.
OpenRouter LLM Rankings
Various
LLM rankings and AI leaderboard based on benchmarks and real usage data from millions of users. See which AI models developers actually use.
Qwen
Various
Qwickly forging AGI, enhancing intelligence.
RunThisLLM
Various
Find out exactly what hardware you need to run any local LLM, image, video, or audio AI model. 275+ models with full build specs and performance estimates.
SEAL LLM Leaderboard
Various
Explore leaderboards with expert-driven LLM benchmarks and updated AI model rankings across coding, reasoning and more.
Vicuna-13B
Various
We introduce Vicuna-13B, an open-source chatbot trained by fine-tuning LLaMA on user-shared conversations collected from ShareGPT. Preliminary evaluation using GPT-4 as a judge s
Whisper
Various
Robust speech recognition via large-scale weak supervision. [#opensource](https://github.com/openai/whisper)
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
