LMDeploy
by Community
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
OSS
LMDeploy
Added 1 June 2026
Overview
LMDeploy is a toolkit for compressing, deploying, and serving large language models. It provides quantization, efficient inference, and a serving backend to reduce model size and latency.
Best for
Best for
Developers who need to compress and serve LLMs efficiently in production
Use cases
- Quantize LLMs to lower precision for faster inference
- Deploy and serve LLMs with a high-performance inference engine
- Integrate LLMs into production pipelines with minimal overhead
Notes
LMDeploy is a toolkit for compressing, deploying, and serving large language models. It provides quantization, efficient inference, and a serving backend to reduce model size and latency.
7,876 stars on GitHub. Last updated 2026-06-01. Licensed Apache-2.0.
Use cases
- Quantize LLMs to lower precision for faster inference
- Deploy and serve LLMs with a high-performance inference engine
- Integrate LLMs into production pipelines with minimal overhead
Pros
- Strong quantization support reduces memory and speeds up inference
- High-performance serving backend with low latency
- Active community with frequent updates and 7.8k GitHub stars
Cons
- Limited to models compatible with its engine and quantization methods
- Documentation and examples may lag behind rapid development
- Requires Python and some familiarity with model deployment tooling
Indexed from awesome-llm and enriched against its public facts.
Pros
- Strong quantization support reduces memory and speeds up inference
- High-performance serving backend with low latency
- Active community with frequent updates and 7.8k GitHub stars
Cons
- Limited to models compatible with its engine and quantization methods
- Documentation and examples may lag behind rapid development
- Requires Python and some familiarity with model deployment tooling
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
vLLM
Community
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang
Community
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT-LLM
Community
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NV
llama.cpp
Community
LLM inference in C/C++
OpenLLM
Community
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
FastChat
Community
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
llama.cpp
Community
LLM inference in C/C++
MNN-LLM
Community
MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.
OpenLLM
Community
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
SGLang
Community
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT-LLM
Community
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NV
vLLM
Community
A high-throughput and memory-efficient inference and serving engine for LLMs
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
LangChain
Community
The agent engineering platform.
Dify
Community
Production-ready platform for agentic workflow development.
Open WebUI
Various
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
