Promptfoo: Test and Red-Team Your LLM Prompts Before Shipping
Promptfoo runs your prompts against multiple models, checks outputs with assertion functions, and red-teams for jailbreaks and PII leakage - all from a YAML config.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
Promptfoo is a CLI tool for testing LLM prompts systematically. You write test cases in YAML, define assertions, and Promptfoo runs every prompt against every model in your config - giving you a comparison matrix and a pass/fail report. It also ships a red-teaming engine that automatically probes for safety vulnerabilities.
Installation
npm install -g promptfoo
# or
npx promptfoo@latest
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Basic YAML Test Config
# promptfooconfig.yaml
providers:
- openai:gpt-4o-mini
- openai:gpt-4o
- ollama:llama3.1:8b
prompts:
- "Summarize the following text in one sentence: {{text}}"
tests:
- vars:
text: "PagedAttention is a memory management technique for LLM inference that treats the KV cache like virtual memory pages."
assert:
- type: contains
value: "KV cache"
- type: javascript
value: "output.length < 200"
- type: llm-rubric
value: "The summary is factually accurate and concise"
Run with:
promptfoo eval
promptfoo view # Open browser UI with comparison table
Model Comparison in PR Checks
Add --ci flag to output JSON results consumable by GitHub Actions:
promptfoo eval --ci --output results.json
Parse results.json in a workflow step to comment score diffs on the PR - your reviewers see exactly which model and which test case regressed.
Custom JavaScript Assertions
For complex validation logic, write a JS function:
// assertions/check-structured.js
module.exports = (output) => {
try {
const parsed = JSON.parse(output);
return {
pass: parsed.name && parsed.age > 0,
score: 1,
reason: "Valid structured output",
};
} catch {
return { pass: false, score: 0, reason: "Not valid JSON" };
}
};
Reference in YAML:
assert:
- type: javascript
value: file://assertions/check-structured.js
Red-Teaming
promptfoo redteam init # Generates redteam config from your system prompt
promptfoo redteam run # Runs attack probes
promptfoo redteam report # View results
Red-team plugins include:
- prompt-injection: attempts to override system prompt
- jailbreak: social engineering and role-play attacks
- pii: probes for PII leakage in RAG responses
- sql-injection: for LLMs with DB tool access
- excessive-agency: checks if the model takes unauthorised actions
GitHub Actions Integration
name: Prompt Quality Check
on: [pull_request]
jobs:
promptfoo:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: npx promptfoo@latest eval --ci
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
Full documentation at promptfoo.dev.

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
GPT-6 Astra Benchmarks, Pricing, and Safety: What Developers Need to Know
OpenAI's GPT-6 Astra delivers state-of-the-art coding and agentic performance, but costs 50% more than GPT-4o. This guide covers benchmarks, pricing, safety, and practical advice for developers deciding whether to upgrade.
How to Test AI Agents with Robust Mocks and Integration Suites
Learn how to test non-deterministic AI agents in CI/CD with contract-based mocks, trajectory evaluation, and statistical Pass^k metrics.
Prompt Versioning and Evaluation in CI/CD Pipelines: A Practical Guide
Treating prompts as code: how to track prompt changes, version them in git, and run automated regression tests on code changes.
// discussion
Comments