How We Benchmark
Get a result you can defend. Every benchmark Metrum AI publishes is held to the same five standards below. Where an older result predates one of these standards, we are updating it; new results are held to all five from first publication.
1. Run count and variance
Every result is reported alongside the number of runs it is based on. Where more than one run was taken, we report the spread (standard deviation or a min/max range), not a single point figure implying more precision than the measurement supports.
2. Warm-up and steady-state criteria
Throughput and latency figures are measured after the system reaches steady state, with the warm-up period excluded and stated. A cold-start figure is only used when the workload we are describing is itself a cold-start scenario, and it is labeled as such.
3. Version pinning
Every result names the serving engine and version (for example vLLM, SGLang, or TensorRT-LLM), the tensor-parallel degree, and the relevant scheduler or quantization settings. Where practical, we pin a specific commit or container digest so the configuration can be reproduced.
4. Platform-confound disclosure
When we compare two pieces of hardware, we state plainly whether the comparison holds every other variable constant, or whether it is a system-level comparison across two different platforms (different server chassis, CPU, memory, or cooling). If the platforms differ, we say so in the same sentence as the headline number, not in a footnote.
5. Synthetic vs. trace-replay workloads
A result built on a fixed synthetic input/output token distribution is labeled synthetic. A result built on a replayed real-world traffic trace is labeled trace-replay. The two are not interchangeable, and we do not present one as the other.
Example in practice: our TCO Titan comparison of MI355X and MI300X measures two different Dell PowerEdge platforms (XE9785L and XE9680), not the accelerators in isolation with every other variable held constant. Read it as a generation-over-generation system comparison, not an accelerator-only benchmark. See the full report.