DeepSpeed
by Community
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
OSS
DeepSpeed
Added 1 June 2026
Overview
DeepSpeed is a Python library for optimizing distributed training and inference of large language models and deep neural networks. It reduces memory footprint, accelerates training speed, and enables efficient multi-GPU and multi-node setups through techniques like gradient checkpointing, mixed precision, and ZeRO optimizer states partitioning.
Best for
Best for
Teams training large models who need to maximize GPU efficiency and scale across multiple devices.
Use cases
- Training large models on limited GPU memory
- Scaling training across multiple GPUs or nodes
- Reducing inference latency for deployed models
Notes
DeepSpeed is a Python library for optimizing distributed training and inference of large language models and deep neural networks. It reduces memory footprint, accelerates training speed, and enables efficient multi-GPU and multi-node setups through techniques like gradient checkpointing, mixed precision, and ZeRO optimizer states partitioning.
42,436 stars on GitHub. Last updated 2026-06-01. Licensed Apache-2.0.
Use cases
- Training large models on limited GPU memory
- Scaling training across multiple GPUs or nodes
- Reducing inference latency for deployed models
Pros
- Significant memory savings enable training larger models on existing hardware
- Production-ready with strong community adoption and Microsoft backing
- Works with existing PyTorch code with minimal integration effort
Cons
- Steep learning curve for advanced features like ZeRO stages and custom configurations
- Debugging distributed training issues remains complex despite optimizations
- Performance gains vary significantly based on hardware, model architecture, and tuning
Indexed from awesome-llm and enriched against its public facts.
Pros
- Significant memory savings enable training larger models on existing hardware
- Production-ready with strong community adoption and Microsoft backing
- Works with existing PyTorch code with minimal integration effort
Cons
- Steep learning curve for advanced features like ZeRO stages and custom configurations
- Debugging distributed training issues remains complex despite optimizations
- Performance gains vary significantly based on hardware, model architecture, and tuning
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
BMTrain
Community
Efficient Training (including pre-training and fine-tuning) for Big Models
Colossal-AI
Community
Making large AI models cheaper, faster and more accessible
Horovod
Community
Distributed training framework for TensorFlow, Keras, PyTorch, and Apache MXNet.
Liger-Kernel
Community
Efficient Triton Kernels for LLM Training
maxtext
Community
A simple, performant and scalable Jax LLM!
Megatron-LM
Community
Ongoing research training transformer models at scale
Mesh Tensorflow
Community
Mesh TensorFlow: Model Parallelism Made Easier
MInference
Community
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to
nanotron
Community
Minimalistic large language model 3D-parallelism training
NeMo Framework
Community
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech
OneComp
Community
Python package for LLM compression
torchtitan
Community
A PyTorch native platform for training generative AI models
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
Accelerate
Community
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP a
Axolotl
Community
Go ahead and axolotl questions
BELLE
Community
BELLE: Be Everyone's Large Language model Engine(开源中文对话大模型)
FastChat
Community
An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
FedML
Community
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables runn
FlagAI
Community
FlagAI (Fast LArge-scale General AI models) is a fast, easy-to-use and extensible toolkit for large-scale model.
GPT-NeoX
Community
An implementation of model parallel autoregressive transformers on GPUs, based on the Megatron and DeepSpeed libraries
Litgpt
Community
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
Llama 3-8|70B
Community
[Llama 2-7 13 70B](https://llama.meta.com/llama2/)
ROLL
Community
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
State of GPT
Community
Go deep on real code and real systems with the teams building and scaling AI at Microsoft Build, June 2–3, 2026, in San Francisco and online.
text-generation-inference
Community
Large Language Model Text Generation Inference
Tune Studio
Community
Playground for devs to finetune & deploy LLMs
veRL
Community
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Together AI
Various
Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research.
BLOOMZ&mT0
Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Flyflow
Community
Open source, high performance fine tuning as a service for GPT4 quality models with 5x lower latency and 3x lower cost
Megatron-DeepSpeed
Community
Ongoing research training transformer language models at scale, including: BERT & GPT-2
Moonlight-A3B
Community
Moonshot's Compute-efficient MoE LLM, first Scaling Up of Muon Optimizer
Bloom
Various
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
DeepSeek
Various
Org profile for DeepSeek on Hugging Face, the AI community building the future.
NeMo
Various
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech
Datatrove
Community
Freeing data processing from scripting madness by providing a set of platform-agnostic customizable pipeline processing blocks.
FastDatasets
Community
A powerful tool for creating high-quality training datasets for Large Language Models (LLMs)(一个快速生成高质量LLM微调训练数据集的工具)
Grok-1-314B-MoE
Community
Grok-1-314B-MoE — indexed from awesome-llm
peft
Community
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
ROLL
Community
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
Slurm
Community
Slurm: A Highly Scalable Workload Manager
Transformer Engine
Community
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide b
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
