Enterprise DNA
Directories / Alternatives / OpenAI Evals

Open Source Alternatives

Open source alternatives to OpenAI Evals

Open source alternatives to OpenAI Evals, ranked by GitHub stars and freshness.

16 open-source alternatives in the index, ranked by GitHub stars and freshness.

O OSS Framework medium

promptfoo

Community

Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative config

★ 21,784 updated 2d ago
open-source

Best for: Teams building LLM applications who need systematic prompt validation and security testing before deployment

O OSS Framework medium

Giskard

Community

🐢 Open-Source Evaluation & Testing library for LLM Agents

★ 5,414 updated 5d ago
open-source

Best for: Python developers building LLM agents who need automated safety and quality testing.

O OSS Framework medium

simple-evals

Community

Eval tools by OpenAI.

★ 4,508 updated 1mo ago
open-source

Best for: Developers who need a straightforward, OpenAI-aligned evaluation toolkit for LLM outputs

O OSS Framework medium

LangWatch

Community

The platform for LLM evaluations and AI agent testing

★ 3,275 updated 2d ago
open-source

Best for: Developers building and testing LLM-based agents in TypeScript who need a lightweight evaluation framework

O OSS Framework medium

HELM

Community

Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducib

★ 2,811 updated 2d ago
open-source

Best for: Researchers and developers who need rigorous, multi-dimensional evaluation of foundation models

O OSS Framework medium

Chain-of-Thought Hub

Community

Benchmarking large language models' complex reasoning ability with chain-of-thought prompting

★ 2,773 updated 1y ago
open-source

Best for: Researchers and developers evaluating LLM reasoning capabilities with chain-of-thought prompting

O OSS Framework medium

instruct-eval

Community

This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.

★ 553 updated 2y ago
open-source

Best for: Researchers and developers who need a simple, standardized way to evaluate instruction-tuned language models

O OSS Framework medium

OLMO-eval

Community

Evaluation suite for LLMs

★ 379 updated 10mo ago
open-source

Best for: Researchers and developers evaluating OLMo or compatible LLMs with reproducible benchmarks

O OSS Framework medium

Berkeley Function-Calling Leaderboard

Community

Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM

open-source

Best for: Developers and researchers evaluating LLMs for tool-use and function-calling applications

O OSS Framework medium

CompassRank

Community

评测榜单旨在为大语言模型和多模态模型提供全面、客观且中立的得分与排名,同时提供多能力维度的评分参考,以便用户能够更全面地了解大模型的能力水平。

open-source

Best for: Developers evaluating and comparing open-source LLMs and multimodal models

O OSS Framework medium

FELM

Community

FELM: Benchmarking Factuality Evaluation of Large Language Models

open-source

Best for: Researchers and developers needing a standardized way to measure LLM factuality

O OSS Framework medium

LawBench

Community

LawBench

open-source

Best for: Researchers and engineers evaluating or selecting LLMs for legal applications

O OSS Framework medium

LLMEval

Community

LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.

open-source

Best for: Researchers and developers building or using LLM evaluation benchmarks

O OSS Framework medium

OlympicArena

Community

OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

open-source

Best for: Researchers and developers evaluating reasoning capabilities of AI models across multiple disciplines.

O OSS Framework medium

SciBench

Community

Evaluating scientific problems

open-source

Best for: Researchers and developers evaluating AI systems on scientific reasoning tasks

O OSS Framework medium

SuperBench

Community

a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on their performance in different aspects such as natural langu

open-source

Best for: Researchers and developers who need a standardized platform to compare LLM performance across common tasks.