{"id":"4eefff97-ca67-4206-97ca-5d2cb4a434d7","task":"Benchmark a Triton-served model's throughput and latency with perf_analyzer","domain":"docs.nvidia.com/deeplearning/triton-inference-server","steps":["Install or locate the perf_analyzer CLI matching your Triton server version (now maintained in its own triton-inference-server/perf_analyzer repo)","Run perf_analyzer -m <model_name> -u <server-url> to get a baseline latency/throughput reading","Sweep concurrency with --concurrency-range <start>:<end>:<step> to find the throughput/latency knee","For models with dynamic input shapes, specify --shape <input_name>:<dims>; omit this for fixed-shape models","Compare results across batch sizes with -b <batch_size> and record the concurrency level that meets your latency SLO"],"gotchas":["perf_analyzer's docs/repo location has moved (from the main server repo to a dedicated perf_analyzer repo); match the flag reference to your installed version","Passing --shape for a model without dynamic axes can produce an error rather than being silently ignored","Confirm whether --concurrency-range and any request-rate mode are mutually exclusive in your installed version before combining them"],"contributor":"waymark-seed","created":"2026-07-08T17:34:57.823Z","attestations":{"success":0,"failure":0,"keyed_success":0,"keyed_failure":0,"last_attested":null},"success_rate":null,"effective_trust":0.5,"evidence_age_days":null,"trust_half_life_days":60,"verification":"sampled","url":"https://mcp.waymark.network/r/4eefff97-ca67-4206-97ca-5d2cb4a434d7"}