Enterprise DNA Enterprise DNA
Directories / Alternatives / OpenAI Evals

Open Source Alternatives

Open source alternatives to OpenAI Evals

Open source alternatives to OpenAI Evals, ranked by GitHub stars and freshness.

24 open-source alternatives in the index, ranked by GitHub stars and freshness.

O OSS Framework medium

promptfoo

Community

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative config

Alternative to Openai Evals, Ragas

★ 21,784 updated 3mo ago
open-source MIT TypeScript

Best for: Teams building LLM applications who need systematic prompt validation and security testing before deployment

O OSS Framework medium

Opik

Community

Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

Alternative to Promptfoo, Ragas, Openai Evals

★ 19,417 updated 3mo ago
open-source Apache-2.0 Python

Best for: Python developers building production LLM systems who need observability and systematic evaluation.

O OSS Framework medium

Ragas

Community

Supercharge Your LLM Application Evaluations 🚀

Alternative to Promptfoo, Openai Evals, Lm Evaluation Harness +1 more

★ 14,186 updated 6mo ago
open-source Apache-2.0 Python

Best for: Teams building RAG systems who need continuous evaluation without manual labeling

O OSS Framework medium

lm-evaluation-harness

Community

A framework for few-shot evaluation of language models.

Alternative to Openai Evals, Promptfoo, Ragas

★ 12,772 updated 4mo ago
open-source MIT Python

Best for: Researchers and engineers benchmarking LLM performance against established academic standards

O OSS Framework medium

Giskard

Community

🐢 Open-Source Evaluation & Testing library for LLM Agents

Alternative to Promptfoo, Ragas, Openai Evals

★ 5,414 updated 3mo ago
open-source Apache-2.0 Python

Best for: Python developers building LLM agents who need automated safety and quality testing.

O OSS Framework medium

simple-evals

Community

Eval tools by OpenAI.

Alternative to Openai Evals, Lm Evaluation Harness, Promptfoo +1 more

★ 4,508 updated 4mo ago
open-source MIT Python

Best for: Developers who need a straightforward, OpenAI-aligned evaluation toolkit for LLM outputs

O OSS Framework medium

Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Community

Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models

Alternative to Openai Evals, Lm Evaluation Harness

★ 3,244 updated 2y ago
open-source Apache-2.0 Python

Best for: Researchers and engineers studying language model capabilities and scaling behavior

O OSS Framework medium

HELM

Community

Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducib

Alternative to Lm Evaluation Harness, Openai Evals

★ 2,811 updated 3mo ago
open-source Apache-2.0 Python

Best for: Researchers and developers who need rigorous, multi-dimensional evaluation of foundation models

O OSS Framework medium

Chain-of-Thought Hub

Community

Benchmarking large language models' complex reasoning ability with chain-of-thought prompting

Alternative to Lm Evaluation Harness, Openai Evals

★ 2,773 updated 2y ago
open-source MIT Jupyter Notebook

Best for: Researchers and developers evaluating LLM reasoning capabilities with chain-of-thought prompting

O OSS Framework medium

lighteval

Community

Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends

Alternative to Lm Evaluation Harness, Openai Evals

★ 2,430 updated 3mo ago
open-source MIT Python

Best for: Researchers and developers who need a unified way to evaluate and compare LLMs from different sources

O OSS Framework medium

instruct-eval

Community

This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.

Alternative to Lm Evaluation Harness, Openai Evals

★ 553 updated 2y ago
open-source Apache-2.0 Python

Best for: Researchers and developers who need a simple, standardized way to evaluate instruction-tuned language models

O OSS Framework medium

OLMO-eval

Community

Evaluation suite for LLMs

Alternative to Lm Evaluation Harness, Openai Evals

★ 379 updated 1y ago
open-source Apache-2.0 Python

Best for: Researchers and developers evaluating OLMo or compatible LLMs with reproducible benchmarks

O OSS Framework medium

AlpacaEval

Community

AlpacaEval Leaderboard

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Researchers and developers benchmarking instruction-tuned language models

O OSS Framework medium

Arize-Phoenix

Community

Arize Phoenix: Open Source AI Development Platform

Alternative to Promptfoo, Opik, Ragas +2 more

open-source

Best for: Developers and teams needing an open-source observability layer for AI applications

O OSS Framework medium

Arthur Shield

Community

Open-source toolkit for building, testing, and monitoring AI agents. Version prompts, run experiments, trace workflows, and catch issues before users do.

Alternative to Promptfoo, Opik, Openai Evals

open-source

Best for: Developers building custom AI agents who need guardrails and observability

O OSS Framework medium

Berkeley Function-Calling Leaderboard

Community

Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Developers and researchers evaluating LLMs for tool-use and function-calling applications

O OSS Framework medium

CompassRank

Community

评测榜单旨在为大语言模型和多模态模型提供全面、客观且中立的得分与排名,同时提供多能力维度的评分参考,以便用户能够更全面地了解大模型的能力水平。

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Developers evaluating and comparing open-source LLMs and multimodal models

O OSS Framework medium

Guardrails.ai

Community

Learn about Guardrails AI and how it helps build reliable AI applications

Alternative to Promptfoo, Openai Evals

open-source

Best for: Developers building production LLM applications that need runtime guardrails for safety, format, and reliability

O OSS Framework medium

InfiBench

Community

IInfiBench: Evaluating the Question-Answering Capabilities of Code LLMs

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Researchers and developers evaluating or comparing code LLMs on question-answering tasks

O OSS Framework medium

LawBench

Community

LawBench

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Researchers and engineers evaluating or selecting LLMs for legal applications

O OSS Framework medium

LLMEval

Community

LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.

Alternative to Lm Evaluation Harness, Openai Evals, Ragas

open-source

Best for: Researchers and developers building or using LLM evaluation benchmarks

O OSS Framework medium

M3CoT

Community

Leaderboard | M 3 CoT

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Researchers and developers evaluating multi-modal chain-of-thought reasoning in AI models

O OSS Framework medium

MathEval

Community

a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.

Alternative to Openai Evals, Lm Evaluation Harness

open-source

Best for: Researchers and developers benchmarking mathematical reasoning in large models.

O OSS Framework medium

MMedBench

Community

Medical Multilingual Benchmark

Alternative to Lm Evaluation Harness, Openai Evals

open-source

Best for: Researchers and developers building multilingual medical AI systems