Open Source Alternatives
Open source alternatives to Ragas
Open source alternatives to Ragas, ranked by GitHub stars and freshness.
10 open-source alternatives in the index, ranked by GitHub stars and freshness.
promptfoo
Community
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative config
Alternative to Openai Evals, Ragas
Best for: Teams building LLM applications who need systematic prompt validation and security testing before deployment
Opik
Community
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Alternative to Promptfoo, Ragas, Openai Evals
Best for: Python developers building production LLM systems who need observability and systematic evaluation.
lm-evaluation-harness
Community
A framework for few-shot evaluation of language models.
Alternative to Openai Evals, Promptfoo, Ragas
Best for: Researchers and engineers benchmarking LLM performance against established academic standards
Evidently
Community
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Alternative to Promptfoo, Ragas
Best for: Data scientists and ML engineers who need a comprehensive, open-source observability framework for both traditional models and LLMs.
Giskard
Community
🐢 Open-Source Evaluation & Testing library for LLM Agents
Alternative to Promptfoo, Ragas, Openai Evals
Best for: Python developers building LLM agents who need automated safety and quality testing.
AutoRAG
Community
AutoRAG: An Open-Source Framework for Retrieval-Augmented Generation (RAG) Evaluation & Optimization with AutoML-Style Automation
Alternative to Ragas
Best for: Developers building and optimizing custom RAG systems for production or research.
simple-evals
Community
Eval tools by OpenAI.
Alternative to Openai Evals, Lm Evaluation Harness, Promptfoo +1 more
Best for: Developers who need a straightforward, OpenAI-aligned evaluation toolkit for LLM outputs
LangWatch
Community
The platform for LLM evaluations and AI agent testing
Alternative to Promptfoo, Opik, Ragas +1 more
Best for: Developers building and testing LLM-based agents in TypeScript who need a lightweight evaluation framework
Arize-Phoenix
Community
Arize Phoenix: Open Source AI Development Platform
Alternative to Promptfoo, Opik, Ragas +2 more
Best for: Developers and teams needing an open-source observability layer for AI applications
LLMEval
Community
LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.
Alternative to Lm Evaluation Harness, Openai Evals, Ragas
Best for: Researchers and developers building or using LLM evaluation benchmarks