Open Source Alternatives
Open source alternatives to LM Eval Harness
Open source alternatives to LM Eval Harness, ranked by GitHub stars and freshness.
20 open-source alternatives in the index, ranked by GitHub stars and freshness.
OpenAI Evals
Community
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Alternative to Lm Evaluation Harness, Promptfoo
Best for: Teams building LLM applications who need systematic, reproducible evaluation workflows
Ragas
Community
Supercharge Your LLM Application Evaluations 🚀
Alternative to Promptfoo, Openai Evals, Lm Evaluation Harness +1 more
Best for: Teams building RAG systems who need continuous evaluation without manual labeling
simple-evals
Community
Eval tools by OpenAI.
Alternative to Openai Evals, Lm Evaluation Harness, Promptfoo +1 more
Best for: Developers who need a straightforward, OpenAI-aligned evaluation toolkit for LLM outputs
LangWatch
Community
The platform for LLM evaluations and AI agent testing
Alternative to Promptfoo, Opik, Ragas +1 more
Best for: Developers building and testing LLM-based agents in TypeScript who need a lightweight evaluation framework
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Community
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models
Alternative to Openai Evals, Lm Evaluation Harness
Best for: Researchers and engineers studying language model capabilities and scaling behavior
HELM
Community
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the Center for Research on Foundation Models (CRFM) at Stanford for holistic, reproducib
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers who need rigorous, multi-dimensional evaluation of foundation models
Chain-of-Thought Hub
Community
Benchmarking large language models' complex reasoning ability with chain-of-thought prompting
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers evaluating LLM reasoning capabilities with chain-of-thought prompting
lighteval
Community
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers who need a unified way to evaluate and compare LLMs from different sources
instruct-eval
Community
This repository contains code to quantitatively evaluate instruction-tuned models such as Alpaca and Flan-T5 on held-out tasks.
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers who need a simple, standardized way to evaluate instruction-tuned language models
OLMO-eval
Community
Evaluation suite for LLMs
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers evaluating OLMo or compatible LLMs with reproducible benchmarks
AlpacaEval
Community
AlpacaEval Leaderboard
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers benchmarking instruction-tuned language models
Arize-Phoenix
Community
Arize Phoenix: Open Source AI Development Platform
Alternative to Promptfoo, Opik, Ragas +2 more
Best for: Developers and teams needing an open-source observability layer for AI applications
Berkeley Function-Calling Leaderboard
Community
Explore The Berkeley Function Calling Leaderboard (also called The Berkeley Tool Calling Leaderboard) to see the LLM
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Developers and researchers evaluating LLMs for tool-use and function-calling applications
CompassRank
Community
评测榜单旨在为大语言模型和多模态模型提供全面、客观且中立的得分与排名,同时提供多能力维度的评分参考,以便用户能够更全面地了解大模型的能力水平。
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Developers evaluating and comparing open-source LLMs and multimodal models
InfiBench
Community
IInfiBench: Evaluating the Question-Answering Capabilities of Code LLMs
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers evaluating or comparing code LLMs on question-answering tasks
LawBench
Community
LawBench
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and engineers evaluating or selecting LLMs for legal applications
LLMEval
Community
LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.
Alternative to Lm Evaluation Harness, Openai Evals, Ragas
Best for: Researchers and developers building or using LLM evaluation benchmarks
M3CoT
Community
Leaderboard | M 3 CoT
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers evaluating multi-modal chain-of-thought reasoning in AI models
MathEval
Community
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.
Alternative to Openai Evals, Lm Evaluation Harness
Best for: Researchers and developers benchmarking mathematical reasoning in large models.
MMedBench
Community
Medical Multilingual Benchmark
Alternative to Lm Evaluation Harness, Openai Evals
Best for: Researchers and developers building multilingual medical AI systems