Skip to main content
Metrum Insights
Live

Throughput vs Concurrency

01.5k3k4.5k6kConcurrent Requests

Output Throughput

MI355X
MI300X
H100 SXM
B200x1

Peak Metrics

Throughput

0.0k tok/s

TTFT

0.0 ms

Concurrency

0+

B200x1
MI355X
H100 SXM
MI300X
v4.1 Live

150+

Models Supported

15+

GPU SKUs

5+

Modalities

30+

Performance Metrics

100+

Benchmark Visualizations

How It Works

From configuration to decision, in three steps.

01

Configure

Define your benchmark matrix: models, chips, precision, concurrency, and serving engines. No code required.

02

Run

Launch parallel benchmark runs across your own environment or ours, real workloads and endpoints, measured end to end.

03

Decide

Filter and compare throughput, latency, and quality (via KYAI) across every run in Pulse, then lock in the configuration that fits your workload.

What's New in v4.0 & v4.1

Latest Features Available

Multimodal benchmarking, Google TPU support, quality evaluation, a shared team workspace, and hardware leaderboards across every performance mode.

New in v4.1

Multi-Node Distributed Inference

Benchmark models too large for one server. A new deployment mode clusters multiple servers with Kubernetes (K3s) and deploys the model via llm-d, with cluster formation and bootstrap fully automated.

Multi-Server Clusters

K3s-orchestrated, llm-d deployed

Zero Manual Setup

Automated cluster formation & bootstrap

Distributed Cluster

Kubernetes-orchestrated via llm-d

Node 1

Node 2

Node 3

Cluster Ready

One model, one benchmark

Automated K3s cluster bootstrap

New in v4.1

KV Cache Offload

Choose where the KV cache lives - GPU, CPU (system RAM), or NVMe, each via a standalone LMCache cache server - so the cost and performance tradeoff can be measured instead of assumed.

GPU / CPU / NVMe

Measured via LMCache

Multi-Turn Chat Bench

Realistic multiroundqa traffic

Cache Tier

Measured via LMCache

GPUFastest, smallest capacity
CPU (RAM)Balanced
NVMeLargest capacity, slower

Reported in Pulse & Excel Export

Benchmarked under multi-turn chat traffic

New in v4.0

Multimodal Benchmarking

LLM, VLM, ASR (Whisper), and ImageGen (vLLM-Omni + SGLang Diffusion) as first-class workload types, benchmarked against built-in reference configurations in a single unified workflow.

LLM & VLM

Text and vision-language models

ASR & ImageGen

Whisper, vLLM-Omni, SGLang Diffusion

Unified Workflow
LLM
VLM
ASR
ImageGen

One Unified Workflow

Reference A
Reference B
Reference C
Baseline

New in v4.0

Google TPU Support

Benchmark on Google TPU v5 and v6 via vLLM, alongside NVIDIA and AMD GPUs, in the same unified workflow.

TPU v5 & v6

Benchmarked via vLLM

Same Workflow

No separate TPU-only tooling

Hardware Coverage

One workflow, every accelerator

NVIDIA GPU

AMD GPU

Google TPU

TPU v5 · TPU v6

Benchmarked via vLLM

New in v4.0

KYAI: Quality Evaluation

KYAI generates a candidate response, then an LLM judge scores it against your rubric, with leaderboards and comparison views across your workspace.

Candidate Generation
LLM Judge Scoring
Public Leaderboard
Comparison Views
Generate & Judge
LIVE

Response

From any model or endpoint

Phase 1: Generate

Candidate evaluation drafted

Phase 2: LLM Judge

Scored against the rubric

Verdict scored
AccuracyCoherenceRelevanceSafety
LeaderboardGPT-4oClaude 3.5Gemini Pro

New in v4.0

Pulse: Team Workspace for Benchmark Runs

A new team workspace that turns raw runs into shared, filterable views. Your whole team can compare results without digging through separate reports.

Team Workspace

Shared views across your org

Filterable Views

Turn raw runs into shared reports

Team Workspace
Run A
Run B
Run C
Run D

Shared View

ModelFrameworkCostHardware
FilterableShared with teamNo digging through reports

New in v4.0

Hardware Leaderboards

Performance rankings across every mode show how today's leading hardware performs across state-of-the-art models, ranked side by side.

Ranked by Mode

Latest chips ranked by performance

Model Comparison

How each chip performs across SOA models

Hardware Leaderboard
ALL MODES
ThroughputLatencyCostEfficiency
#1

B300

Llama 4

#2

B200

Qwen 3.6

#3

MI355X

DeepSeek V4

#4

TPU v6

GLM 5.2

Llama 4

Qwen 3.6

DeepSeek V4

GLM 5.2

Fully Automated AI Performance Benchmarking

Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.

Workloads

Purpose-built tools for every benchmark workload.

Quality Evaluation (KYAI) WorkloadGenAI-Perf WorkloadInferenceX WorkloadLarge Language Model WorkloadVision-Language Model WorkloadAudio WorkloadImage Generation Workload

Frameworks & Configurations

Configure and test across multiple dimensions simultaneously.

Frameworks

vLLMSGLangTensorRT-LLMand more coming soon

Configurations

ConcurrencyISL / OSLTensor ParallelPrecisionQuantizationKV Cache DtypeReasoning ParserStreaming (On/Off)Inference Mode (Chat/Completion)KV Cache OffloadMultinode InferenceGPU CountData Paralleland many more

Chips & Models

Benchmark across architectures and foundation models.

Chips

NVIDIA DatacenterAMD InstinctGoogle TPUIntel Gaudi 3AMD EPYCIntel XeonRTX GPUs

Models

Gemma 4MiniMax M3Kimi K2.7NVIDIA Nemotron 3GLM 5.2DeepSeek V4Qwen 3.6GPT OSSLlama 4and many more

Real-Time Metrics

Track performance across every dimension.

Performance Metrics

Token Throughput

tok/s

Request Throughput

req/s

TTFT

P50 · P90 · P99

TPOT

P50 · P90 · P99

System Power

Watts

GPU Power

Watts

CPU Power

Watts

Efficiency

tok/s/W

Error Rate

%

System Metrics

GPU Utilization

%

GPU Memory

GB

GPU Temperature

°C

CPU Usage

%

CPU Frequency

GHz

Disk I/O

MB/s

SSD Metrics

IOPS

Memory Usage

GB

Where Metrum Insights Fits

Public leaderboards, open-source GPU benchmark suites, and industry standards solve adjacent problems. Here is where each sits.

Compare Metrum vs.

Supports customer-owned infrastructure

Metrum InsightsFull support
Artificial AnalysisNone

Multi-framework single-node benchmarking (vLLM, SGLang, TensorRT-LLM)

Metrum InsightsFull support
Artificial AnalysisNone

Multi-hardware coverage (NVIDIA, AMD, Google TPU) in one platform

Metrum InsightsFull support
Artificial AnalysisPartial

Built-in accuracy and quality evaluation

Metrum InsightsFull support
Artificial AnalysisFull support

Built-in third-party benchmark tool support

Metrum InsightsFull support
Artificial AnalysisNone

Multi-modality benchmarking (LLM, VLM, ASR, ImageGen, agentic) in one platform

Metrum InsightsFull support
Artificial AnalysisNone

Based on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.

Pricing

Your infrastructure.
Our benchmarking platform.

Benchmarks run against your own cloud instances, on-premises servers, and inference endpoints. Your infrastructure bill goes directly to your provider.

Bring your own computeNo infrastructure markupCancel any time

ASR, VLM, and image generation

Multimodal

Starting from

$30,000/yr
ASR, VLM, and image generation workloads
45 targets included
Dedicated Solutions Engineer
Contact sales

Full platform with power tools

Pro

Starting from

$90,000/yr
All Power Tools included
75 targets and Open-Weight Leaderboard
Dedicated CSM and Solutions Engineer
Contact sales

On-premises and air-gapped deployments

Enterprise

Starting from

$180,000/yr
On-premises self-managed install
Unlimited targets under fair use
Dedicated TAM and SE
Contact sales

Questions