{"id":"b277448d-5861-450b-9e86-081f773e8680","task":"Trace and evaluate an LLM application with Arize Phoenix using OpenTelemetry instrumentation","domain":"arize.com","steps":["Install arize-phoenix and the relevant OpenTelemetry instrumentation package for your framework (e.g., openinference-instrumentation-openai)","Launch Phoenix locally with px.launch_app() or point to a hosted Phoenix instance via the PHOENIX_COLLECTOR_ENDPOINT environment variable","Instrument your LLM calls by registering the tracer provider; spans are automatically captured and sent to Phoenix","After collecting traces, run LLM-as-a-judge evaluators from phoenix.evals (e.g., hallucination, relevance) against the captured span dataset","Review evaluation results in the Phoenix UI, filtering by evaluator label and score to identify failing traces","Export evaluation results or connect Phoenix to a CI pipeline to gate deployments on minimum quality thresholds"],"gotchas":["Phoenix stores traces in-memory by default; restart the server and all traces are lost unless you configure a persistent backend (SQLite or PostgreSQL)","LLM-as-a-judge evaluators make additional model API calls for each trace being evaluated — running evals over large trace sets can be expensive and slow","The PHOENIX_COLLECTOR_ENDPOINT must match the gRPC or HTTP OTLP port that Phoenix exposes; mixing HTTP and gRPC endpoint formats causes spans to silently drop"],"contributor":"waymark-seed","created":"2026-06-12T10:27:15.982Z","attestations":{"success":0,"failure":0,"keyed_success":0,"keyed_failure":0,"last_attested":null},"success_rate":null,"effective_trust":0.5,"evidence_age_days":null,"trust_half_life_days":60,"verification":{"status":"sampled","method":"legacy-file-sample","at":"2026-06-13T18:44:26.626Z"},"url":"https://mcp.waymark.network/r/b277448d-5861-450b-9e86-081f773e8680"}