promptfoo
by Community
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative config
OSS
promptfoo
Added 1 June 2026
Overview
promptfoo is a testing framework for evaluating prompts, agents, and RAG systems across multiple LLM providers including GPT, Claude, Gemini, and DeepSeek. It runs comparative benchmarks, red team tests, and vulnerability scans using declarative YAML configs with CLI and CI/CD support.
Best for
Best for
Teams building LLM applications who need systematic prompt validation and security testing before deployment
Use cases
- Compare prompt performance across different LLM models before production
- Automate security testing and adversarial input scanning for AI applications
- Integrate prompt evaluation into CI/CD pipelines for continuous quality checks
Notes
promptfoo is a testing framework for evaluating prompts, agents, and RAG systems across multiple LLM providers including GPT, Claude, Gemini, and DeepSeek. It runs comparative benchmarks, red team tests, and vulnerability scans using declarative YAML configs with CLI and CI/CD support.
21,784 stars on GitHub. Last updated 2026-06-01. Licensed MIT.
Use cases
- Compare prompt performance across different LLM models before production
- Automate security testing and adversarial input scanning for AI applications
- Integrate prompt evaluation into CI/CD pipelines for continuous quality checks
Pros
- Multi-model comparison built in, reducing vendor lock-in risk
- Red teaming and vulnerability scanning included, not bolted on
- Declarative config approach makes tests reproducible and version-controllable
Cons
- Requires familiarity with YAML config syntax and CLI tooling
- Testing scope limited to prompt and agent behavior, not full application integration
- Costs scale with API calls to external LLM providers during test runs
Indexed from awesome-llm and enriched against its public facts.
Pros
- Multi-model comparison built in, reducing vendor lock-in risk
- Red teaming and vulnerability scanning included, not bolted on
- Declarative config approach makes tests reproducible and version-controllable
Cons
- Requires familiarity with YAML config syntax and CLI tooling
- Testing scope limited to prompt and agent behavior, not full application integration
- Costs scale with API calls to external LLM providers during test runs
Open-source & AI alternatives
Swap-in tools that solve the same job. Weigh the trade-offs before you commit.
Arize-Phoenix
Community
Arize Phoenix: Open Source AI Development Platform
Arthur Shield
Community
Open-source toolkit for building, testing, and monitoring AI agents. Version prompts, run experiments, trace workflows, and catch issues before users do.
Evidently
Community
Evidently is an open-source ML and LLM observability framework. Evaluate, test, and monitor any AI-powered system or data pipeline. From tabular data to Gen AI. 100+ metrics.
Giskard
Community
🐢 Open-Source Evaluation & Testing library for LLM Agents
Guardrails.ai
Community
Learn about Guardrails AI and how it helps build reliable AI applications
LangWatch
Community
The platform for LLM evaluations and AI agent testing
lm-evaluation-harness
Community
A framework for few-shot evaluation of language models.
OpenAI Evals
Community
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Opik
Community
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Promptify
Community
Prompt Engineering | Prompt Versioning | Use GPT or other prompt based models to get structured output. Join our discord for Prompt-Engineering, LLMs and other latest research
PromptPerfect
Community
PromptPerfect - AI Prompt Generator and Optimizer
Ragas
Community
Supercharge Your LLM Application Evaluations 🚀
simple-evals
Community
Eval tools by OpenAI.
Pairs with
Other entries in the index that connect to this one. Click through to see the chain.
LangChain
Community
The agent engineering platform.
Opik
Community
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
AdalFlow
Community
AdalFlow: The library to build & auto-optimize LLM applications.
awesome-hallucination-detection
Community
List of papers on hallucination detection in LLMs.
Awesome-LLM-hallucination
Community
LLM hallucination paper list
Awesome LLM Security
Community
A curation of awesome tools, documents and projects about LLM Security.
DSPy
Stanford NLP
Programming, not prompting. Declare what you want, compile prompts and weights against an objective.
FELM
Community
FELM: Benchmarking Factuality Evaluation of Large Language Models
IntelliServer
Community
AI models as scalable microservices, enabling evaluation of LLMs and offering end-to-end functions such as chatbot, semantic search, image generation and beyond.
OpenAI o3-mini
Community
Pushing the frontier of cost-effective reasoning.
Prompt Engineering
Community
Prompt Engineering, also known as In-Context Prompting, refers to methods for how to communicate with LLM to steer its behavior for desired outcomes without updating the model we
Scale Spellbook
Community
Accelerate and scale Generative AI across your enterprise with the platform to transform your data into customized enterprise-ready Generative AI applications.
Get the free Developer’s Field Guide
A 27-page field guide to the AI coding workflow with Claude. Claude Code, MCP servers, the prompt patterns that work, and what to delegate. Free.
Enter your work email. We send it straight over, plus a few short notes worth knowing. Unsubscribe any time.
