DVC
by Community
š¦ Data Versioning and ML Experiments
OSS
DVC
Added 1 June 2026
Overview
DVC (Data Version Control) is a version control system for machine learning projects that tracks data, models, and experiment metadata alongside code. It integrates with Git to manage large files and pipelines, enabling reproducible ML workflows without storing binaries in repositories.
Best for
Best for
ML teams building reproducible pipelines who need Git-like versioning for data and models
Use cases
- Track dataset versions and model artifacts across experiment iterations
- Reproduce ML pipelines and results from previous runs
- Collaborate on ML projects with versioned data and experiment history
Notes
DVC (Data Version Control) is a version control system for machine learning projects that tracks data, models, and experiment metadata alongside code. It integrates with Git to manage large files and pipelines, enabling reproducible ML workflows without storing binaries in repositories.
15,643 stars on GitHub. Last updated 2026-06-01. Licensed Apache-2.0.
Use cases
- Track dataset versions and model artifacts across experiment iterations
- Reproduce ML pipelines and results from previous runs
- Collaborate on ML projects with versioned data and experiment history
Pros
- Integrates seamlessly with Git for unified project versioning
- Handles large files and remote storage without bloating repositories
- Tracks full experiment lineage including parameters, metrics, and outputs
Cons
- Requires Python and command-line familiarity for typical workflows
- Learning curve for teams unfamiliar with version control concepts
- Remote storage setup and configuration adds operational overhead
Indexed from awesome-llmops and enriched against its public facts.
Pros
- Integrates seamlessly with Git for unified project versioning
- Handles large files and remote storage without bloating repositories
- Tracks full experiment lineage including parameters, metrics, and outputs
Cons
- Requires Python and command-line familiarity for typical workflows
- Learning curve for teams unfamiliar with version control concepts
- Remote storage setup and configuration adds operational overhead
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
Airflow
Community
Platform created by the community to programmatically author, schedule and monitor workflows.
ArtiVC
Community
A version control system to manage large files.
Hamilton
Community
Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere pytho
LakeFS
Community
lakeFS - Data version control for your data lake | Git for data
Metaflow
Community
Build, Manage and Deploy AI/ML Systems
Pachyderm
Community
Data-Centric Pipelines and Data Versioning
Ploomber
Community
The fastest ā”ļø way to build data pipelines. Develop iteratively, deploy anywhere. āļø
Quilt
Community
Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
PyTorch
Community
Tensors and Dynamic neural networks in Python with strong GPU acceleration
TensorFlow
Community
An Open Source Machine Learning Framework for Everyone
scikit-learn
Community
scikit-learn: machine learning in Python
Aim
Community
Aim š« ā An easy-to-use & supercharged open-source experiment tracker.
Argo Workflows
Community
Workflow Engine for Kubernetes
Awesome Open MLOps
Community
The Fuzzy Labs guide to the universe of open source MLOps
Awesome Production Machine Learning
Community
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
Delta-Lake
Community
An open-source storage framework that enables building a Lakehouse architecture with compute engines including Spark, PrestoDB, Flink, Trino, and Hive and APIs
Dolt
Community
Dolt ā Git for Data
Great Expectations
Community
Always know what to expect from your data.
Guild AI
Community
Experiment tracking, ML developer tools
JuiceFS
Community
JuiceFS is a distributed POSIX file system built on top of Redis and S3.
Kedro-Viz
Community
Visualise your Kedro data and machine-learning pipelines and track your experiments.
Kedro
Community
Kedro is a toolbox for production-ready data science. It uses software engineering best practices to help you create data engineering and data science pipelines that are reproducib
MLEM
Community
š¶ A tool to package, serve, and deploy any ML model on any platform. Archived to be resurrected one dayš¤
MLRun
Community
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environ
Model Search
Community

Prefect
Community
Prefect is a workflow orchestration framework for building resilient data pipelines in Python.
Starwhale
Community
an MLOps/LLMOps platform
Weco Observe
Community
Build and Optimize your machine learning pipeline with the Weco Platform - based on AIDE ML, the LLM-powered code optimization Agent for Machine Learning Engineering.
Get the free Developerās Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
