Build a promptfoo eval config that tests prompts across multiple providers with assertions, then run and review results

domain: promptfoo.dev · 5 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Run `promptfoo init` (or `init --example getting-started`) to scaffold a `promptfooconfig.yaml`
  2. Add `prompts` with `{{variable}}` placeholders and list target `providers` (e.g. openai:chat:<model>, anthropic:messages:<model>)
  3. Add `tests` with `vars` for each case and optional `assert` entries (e.g. `contains`, `llm-rubric`, `cost`, `latency`)
  4. Run `promptfoo eval` to execute every prompt/provider/test-case combination
  5. Run `promptfoo view` to open the web viewer and compare outputs side by side

Known gotchas

Related routes

Build a promptfoo eval config to regression-test outputs across model versions
promptfoo.dev · 5 steps · unrated
Gate CI on LLM evals with promptfoo
promptfoo.dev · 6 steps · unrated
Gate CI pipeline deployments on LLM eval pass rates using promptfoo
www.promptfoo.dev · 6 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans