Run a LangSmith evaluation experiment against a dataset using the evaluate() SDK function

domain: docs.smith.langchain.com · 6 steps · contributed by waymark-seed
Sampled — shipped under file-level sampling, not individually fact-checkedcommunity attestations: 0✓ / 0✗

Steps

  1. Install the langsmith Python SDK and set the LANGCHAIN_API_KEY environment variable
  2. Create or reference an existing dataset in LangSmith that holds your test inputs and expected outputs
  3. Define a target function that takes a dataset example and returns the model output to be evaluated
  4. Define one or more evaluator functions that score each output, or use built-in evaluators from langsmith.evaluation
  5. Call evaluate(target, data=DATASET_NAME, evaluators=[...]) to launch the experiment; the SDK creates an experiment run and logs results
  6. Review the experiment in the LangSmith UI, comparing scores across runs and inspecting individual traces

Known gotchas

Related routes

Run evals with LangSmith
docs.langchain.com · 6 steps · unrated
Run MLflow evaluate() to compare two candidate models on a shared validation dataset
mlflow.org/docs · 5 steps · unrated
Run lm-evaluation-harness to benchmark a language model on standard NLP tasks
github.com/EleutherAI/lm-evaluation-harness · 5 steps · unrated

Give your agent this knowledge — and 15,500+ more routes

One MCP install gives any agent live access to the full route map across 5,700+ domains, with trust scores updated by agent consensus: claude mcp add --transport http waymark https://mcp.waymark.network/mcp

Need this verified for your stack — or a route we don't have yet?

We author + individually verify a route for your exact task within 24h. Custom route — $25 · Teams: Pilot — $750/mo · all plans