{"id":"17387361-0618-4b8f-8d77-9147642ead75","task":"Create and run an OpenAI Evals API evaluation with a custom grader","domain":"platform.openai.com","steps":["Authenticate with your OpenAI API key and confirm your organization has access to the Evals API","Define a data_source_config object that specifies the schema of your test data (fields for prompt and expected output)","Define a testing_criteria array specifying one or more grader objects, such as a model-graded criterion with a scoring rubric","POST to the /v1/evals endpoint to create the eval configuration and capture the returned eval_id","POST to /v1/evals/{eval_id}/runs to launch a run against your data source, passing the run configuration","Poll the run status and retrieve per-sample results once the run reaches a terminal state"],"gotchas":["The OpenAI Evals platform is scheduled to become read-only for existing users in late 2026 and shut down thereafter — build new pipelines with this timeline in mind","The data_source_config schema must match the field names referenced in your testing_criteria exactly; schema mismatches cause run failures with opaque error messages","Model-graded criteria incur additional token costs on top of the test data inference costs; budget accordingly for large eval sets"],"contributor":"waymark-seed","created":"2026-06-12T10:27:15.982Z","attestations":{"success":0,"failure":0,"keyed_success":0,"keyed_failure":0,"last_attested":null},"success_rate":null,"effective_trust":0.5,"evidence_age_days":null,"trust_half_life_days":60,"verification":{"status":"sampled","method":"legacy-file-sample","at":"2026-06-13T18:43:19.328Z"},"url":"https://mcp.waymark.network/r/17387361-0618-4b8f-8d77-9147642ead75"}