Benchmark any model on any chip.Get a result you can defend.
Independent, reproducible, audit-ready
Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.
Throughput vs Concurrency
Output Throughput
Peak Metrics
Throughput
0.0k tok/s
TTFT
0.0 ms
Concurrency
0+
150+
Models Supported
15+
GPU SKUs
5+
Modalities
30+
Performance Metrics
100+
Benchmark Visualizations
How It Works
From configuration to decision, in three steps.
Configure
Define your benchmark matrix: models, chips, precision, concurrency, and serving engines. No code required.
Run
Launch parallel benchmark runs across your own environment or ours, real workloads and endpoints, measured end to end.
Decide
Filter and compare throughput, latency, and quality (via KYAI) across every run in Pulse, then lock in the configuration that fits your workload.
What's New in v4.0 & v4.1
Latest Features Available
Multimodal benchmarking, Google TPU support, quality evaluation, a shared team workspace, and hardware leaderboards across every performance mode.
New in v4.1
Multi-Node Distributed Inference
Benchmark models too large for one server. A new deployment mode clusters multiple servers with Kubernetes (K3s) and deploys the model via llm-d, with cluster formation and bootstrap fully automated.
Multi-Server Clusters
K3s-orchestrated, llm-d deployed
Zero Manual Setup
Automated cluster formation & bootstrap
Distributed Cluster
Kubernetes-orchestrated via llm-d
Node 1
Node 2
Node 3
One model, one benchmark
Automated K3s cluster bootstrap
New in v4.1
KV Cache Offload
Choose where the KV cache lives - GPU, CPU (system RAM), or NVMe, each via a standalone LMCache cache server - so the cost and performance tradeoff can be measured instead of assumed.
GPU / CPU / NVMe
Measured via LMCache
Multi-Turn Chat Bench
Realistic multiroundqa traffic
Cache Tier
Measured via LMCache
Reported in Pulse & Excel Export
Benchmarked under multi-turn chat traffic
New in v4.0
Multimodal Benchmarking
LLM, VLM, ASR (Whisper), and ImageGen (vLLM-Omni + SGLang Diffusion) as first-class workload types, benchmarked against built-in reference configurations in a single unified workflow.
LLM & VLM
Text and vision-language models
ASR & ImageGen
Whisper, vLLM-Omni, SGLang Diffusion
One Unified Workflow
New in v4.0
Google TPU Support
Benchmark on Google TPU v5 and v6 via vLLM, alongside NVIDIA and AMD GPUs, in the same unified workflow.
TPU v5 & v6
Benchmarked via vLLM
Same Workflow
No separate TPU-only tooling
Hardware Coverage
One workflow, every accelerator
NVIDIA GPU
AMD GPU
Google TPU
TPU v5 · TPU v6
Benchmarked via vLLM
New in v4.0
KYAI: Quality Evaluation
KYAI generates a candidate response, then an LLM judge scores it against your rubric, with leaderboards and comparison views across your workspace.
Response
From any model or endpoint
Phase 1: Generate
Candidate evaluation drafted
Phase 2: LLM Judge
Scored against the rubric
New in v4.0
Pulse: Team Workspace for Benchmark Runs
A new team workspace that turns raw runs into shared, filterable views. Your whole team can compare results without digging through separate reports.
Team Workspace
Shared views across your org
Filterable Views
Turn raw runs into shared reports
Shared View
New in v4.0
Hardware Leaderboards
Performance rankings across every mode show how today's leading hardware performs across state-of-the-art models, ranked side by side.
Ranked by Mode
Latest chips ranked by performance
Model Comparison
How each chip performs across SOA models
B300
Llama 4
B200
Qwen 3.6
MI355X
DeepSeek V4
TPU v6
GLM 5.2
Llama 4
Qwen 3.6
DeepSeek V4
GLM 5.2
Fully Automated AI Performance Benchmarking
Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.
Workloads
Purpose-built tools for every benchmark workload.
Frameworks & Configurations
Configure and test across multiple dimensions simultaneously.
Frameworks
Configurations
Chips & Models
Benchmark across architectures and foundation models.
Chips
Models
Real-Time Metrics
Track performance across every dimension.
Performance Metrics
Token Throughput
tok/s
Request Throughput
req/s
TTFT
P50 · P90 · P99
TPOT
P50 · P90 · P99
System Power
Watts
GPU Power
Watts
CPU Power
Watts
Efficiency
tok/s/W
Error Rate
%
System Metrics
GPU Utilization
%
GPU Memory
GB
GPU Temperature
°C
CPU Usage
%
CPU Frequency
GHz
Disk I/O
MB/s
SSD Metrics
IOPS
Memory Usage
GB
Where Metrum Insights Fits
Public leaderboards, open-source GPU benchmark suites, and industry standards solve adjacent problems. Here is where each sits.
| Capability | ![]() | Artificial Analysis API benchmarking | SemiAnalysis InferenceMAX | MLPerf MLCommons suite | In-house scripts DIY harness |
|---|---|---|---|---|---|
| Supports customer-owned infrastructure | |||||
| Multi-framework single-node benchmarking (vLLM, SGLang, TensorRT-LLM) | |||||
| Multi-hardware coverage (NVIDIA, AMD, Google TPU) in one platform | |||||
| Built-in accuracy and quality evaluation | |||||
| Built-in third-party benchmark tool support | |||||
| Multi-modality benchmarking (LLM, VLM, ASR, ImageGen, agentic) in one platform |
Compare Metrum vs.
Supports customer-owned infrastructure
Full supportMulti-framework single-node benchmarking (vLLM, SGLang, TensorRT-LLM)
Full supportMulti-hardware coverage (NVIDIA, AMD, Google TPU) in one platform
Full supportBuilt-in accuracy and quality evaluation
Full supportBuilt-in third-party benchmark tool support
Full supportMulti-modality benchmarking (LLM, VLM, ASR, ImageGen, agentic) in one platform
Full supportBased on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.
Your infrastructure.
Our benchmarking platform.
Benchmarks run against your own cloud instances, on-premises servers, and inference endpoints. Your infrastructure bill goes directly to your provider.
ASR, VLM, and image generation
Multimodal
Starting from
Full platform with power tools
Pro
Starting from
On-premises and air-gapped deployments
Enterprise
Starting from
Questions
Ready to Accelerate Your AI Performance?
Join industry leaders using Metrum Insights to optimize their AI infrastructure.