Open Source Alternatives
Open source alternatives to Promptfoo
Open source alternatives to Promptfoo, ranked by GitHub stars and freshness.
13 open-source alternatives in the index, ranked by GitHub stars and freshness.
Opik
Community
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Alternative to Promptfoo, Ragas, Openai Evals
Best for: Python developers building production LLM systems who need observability and systematic evaluation.
OpenAI Evals
Community
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Alternative to Lm Evaluation Harness, Promptfoo
Best for: Teams building LLM applications who need systematic, reproducible evaluation workflows
Ragas
Community
Supercharge Your LLM Application Evaluations 🚀
Alternative to Promptfoo, Openai Evals, Lm Evaluation Harness +1 more
Best for: Teams building RAG systems who need continuous evaluation without manual labeling
lm-evaluation-harness
Community
A framework for few-shot evaluation of language models.
Alternative to Openai Evals, Promptfoo, Ragas
Best for: Researchers and engineers benchmarking LLM performance against established academic standards
Evidently
Community
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Alternative to Promptfoo, Ragas
Best for: Data scientists and ML engineers who need a comprehensive, open-source observability framework for both traditional models and LLMs.
Giskard
Community
🐢 Open-Source Evaluation & Testing library for LLM Agents
Alternative to Promptfoo, Ragas, Openai Evals
Best for: Python developers building LLM agents who need automated safety and quality testing.
Promptify
Community
Prompt Engineering | Prompt Versioning | Use GPT or other prompt based models to get structured output. Join our discord for Prompt-Engineering, LLMs and other latest research
Alternative to Guidance, Outlines, Promptfoo
Best for: Python developers seeking a straightforward way to produce structured outputs from LLM prompts while managing prompt versions.
simple-evals
Community
Eval tools by OpenAI.
Alternative to Openai Evals, Lm Evaluation Harness, Promptfoo +1 more
Best for: Developers who need a straightforward, OpenAI-aligned evaluation toolkit for LLM outputs
LangWatch
Community
The platform for LLM evaluations and AI agent testing
Alternative to Promptfoo, Opik, Ragas +1 more
Best for: Developers building and testing LLM-based agents in TypeScript who need a lightweight evaluation framework
Arize-Phoenix
Community
Arize Phoenix: Open Source AI Development Platform
Alternative to Promptfoo, Opik, Ragas +2 more
Best for: Developers and teams needing an open-source observability layer for AI applications
Arthur Shield
Community
Open-source toolkit for building, testing, and monitoring AI agents. Version prompts, run experiments, trace workflows, and catch issues before users do.
Alternative to Promptfoo, Opik, Openai Evals
Best for: Developers building custom AI agents who need guardrails and observability
Guardrails.ai
Community
Learn about Guardrails AI and how it helps build reliable AI applications
Alternative to Promptfoo, Openai Evals
Best for: Developers building production LLM applications that need runtime guardrails for safety, format, and reliability
PromptPerfect
Community
PromptPerfect - AI Prompt Generator and Optimizer
Alternative to Promptfoo, Gpt Prompt Engineer
Best for: Developers and power users who frequently interact with LLMs and want to improve prompt reliability