Enterprise DNA Enterprise DNA
Directories / Alternatives / TensorRT-LLM

Open Source Alternatives

Open source alternatives to TensorRT-LLM

Open source alternatives to TensorRT-LLM, ranked by GitHub stars and freshness.

8 open-source alternatives in the index, ranked by GitHub stars and freshness.

O OSS Framework medium

llama.cpp

Community

LLM inference in C/C++

Alternative to Vllm, Sglang, Tensorrt Llm +2 more

★ 114,160 updated 3mo ago
open-source MIT C++

Best for: Developers building privacy-first or offline-capable applications with constrained hardware

O OSS Framework medium

vLLM

Community

A high-throughput and memory-efficient inference and serving engine for LLMs

Alternative to Tensorrt Llm, Lmdeploy, Sglang +1 more

★ 81,619 updated 3mo ago
open-source Apache-2.0 Python

Best for: Teams building production LLM APIs and services that need to maximize throughput and minimize latency under concurrent load.

O OSS Framework medium

SGLang

Community

SGLang is a high-performance serving framework for large language models and multimodal models.

Alternative to Vllm, Tensorrt Llm, Lmdeploy

★ 28,885 updated 3mo ago
open-source Apache-2.0 Python

Best for: Teams building production LLM services who need performance-optimized serving infrastructure

O OSS Framework medium

MNN-LLM

Community

MNN: A blazing-fast, lightweight inference engine battle-tested by Alibaba, powering high-performance on-device LLMs and Edge AI.

Alternative to Llama Cpp, Tensorrt Llm, Lmdeploy

★ 15,353 updated 3mo ago
open-source Apache-2.0 C++

Best for: Developers building production on-device LLM and edge AI applications where latency and resource efficiency are critical.

O OSS Framework medium

LMDeploy

Community

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Alternative to Vllm, Sglang, Tensorrt Llm +2 more

★ 7,876 updated 3mo ago
open-source Apache-2.0 Python

Best for: Developers who need to compress and serve LLMs efficiently in production

O OSS Framework medium

mistral.rs

Community

Fast, flexible LLM inference

Alternative to Llama Cpp, Vllm, Sglang +2 more

★ 7,205 updated 3mo ago
open-source MIT Rust

Best for: Rust developers seeking a fast, flexible LLM inference framework for performance-critical or resource-constrained environments.

O OSS Framework medium

FasterTransformer

Community

Transformer related optimization, including BERT, GPT

Alternative to Vllm, Tensorrt Llm

★ 6,418 updated 2y ago
open-source Apache-2.0 C++

Best for: Developers seeking maximum inference performance for transformer models on NVIDIA hardware

O OSS Framework medium

MInference

Community

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to

Alternative to Vllm, Sglang, Tensorrt Llm +1 more

★ 1,217 updated 5mo ago
open-source MIT Python

Best for: Developers optimizing long-context LLM inference on NVIDIA GPUs