XGBoost
by Community
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C++ and more. Runs on single machine, Hadoop, Spark, Dask, Flink and D
OSS
XGBoost
Added 1 June 2026
Overview
XGBoost is a gradient boosting library that trains decision tree ensembles for classification, regression, and ranking tasks. It runs on single machines or distributed systems like Spark, Hadoop, and Dask, with bindings for Python, R, Java, Scala, and C++.
Best for
Best for
Data scientists and ML engineers building production models on structured datasets.
Use cases
- Building high-accuracy predictive models for tabular data
- Training models at scale across distributed clusters
- Competing in machine learning competitions
Notes
XGBoost is a gradient boosting library that trains decision tree ensembles for classification, regression, and ranking tasks. It runs on single machines or distributed systems like Spark, Hadoop, and Dask, with bindings for Python, R, Java, Scala, and C++.
28,431 stars on GitHub. Last updated 2026-05-28. Licensed Apache-2.0.
Use cases
- Building high-accuracy predictive models for tabular data
- Training models at scale across distributed clusters
- Competing in machine learning competitions
Pros
- Consistently outperforms other gradient boosting implementations on structured data
- Handles both single-machine and distributed training without code changes
- Mature ecosystem with extensive documentation and community support
Cons
- Requires careful hyperparameter tuning to avoid overfitting
- Slower than simpler models for real-time inference on resource-constrained devices
- Works best on tabular data, not designed for images or text
Indexed from awesome-llmops and enriched against its public facts.
Pros
- Consistently outperforms other gradient boosting implementations on structured data
- Handles both single-machine and distributed training without code changes
- Mature ecosystem with extensive documentation and community support
Cons
- Requires careful hyperparameter tuning to avoid overfitting
- Slower than simpler models for real-time inference on resource-constrained devices
- Works best on tabular data, not designed for images or text
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
scikit-learn
Community
scikit-learn: machine learning in Python
LightGBM
Community
A fast, distributed, high performance gradient boosting (GBT, GBDT, GBRT, GBM or MART) framework based on decision tree algorithms, used for ranking, classification and many other
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
Deepchecks
Community
Deepchecks: Tests for Continuous Validation of ML Models & Data. Deepchecks is a holistic open-source solution for all of your AI & ML validation needs, enabling to thoroughly test
EvalML
Community
EvalML is an AutoML library written in python.
FEDOT
Community
Automated modeling and machine learning framework FEDOT
FLAML
Community
A fast library for AutoML and tuning. Join our Discord: https://discord.gg/Cppx2vSPVP.
Pycaret
Community
Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.
TPOT
Community
The Tree-Based Pipeline Optimization Tool (TPOT) was one of the very first AutoML methods and open-source software packages developed for the data science community. TPOT was dev
dtreeviz
Community
A python library for decision tree visualization and model interpretation.
Feast
Community
The Open Source Feature Store for AI/ML
FeatureTools
Community
An open source python library for automated feature engineering
Goptuna
Community
A hyperparameter optimization framework, inspired by Optuna.
HpBandSter
Community
a distributed Hyperband implementation on Steroids
Hyperband
Community
Tuning hyperparams fast with Hyperband
Hyperopt
Community
Distributed Asynchronous Hyperparameter Optimization in Python
hyperunity
Community
A toolset for black-box hyperparameter optimisation.
Metaflow
Community
Build, Manage and Deploy AI/ML Systems
NNI
Community
An open source AutoML toolkit for automate machine learning lifecycle, including feature engineering, neural architecture search, model compression and hyper-parameter tuning.
Spearmint
Community
Spearmint Bayesian optimization codebase
Upgini
Community
Data search & enrichment library for Machine Learning → Easily find and add relevant features to your ML & AI pipeline from hundreds of public and premium external data sources, in
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
