Skip to main content
Metrum Insights v3.9 is live.

Latest Features Available

VectorDB benchmarking, KV cache SSD offloading, and upgraded runtimes with support for the latest models.

VectorDB Bench (Milvus)

End-to-end vector database benchmarking with HNSW and DISKANN indexing, plus new storage metrics including IOPS, latency, throughput, queue depth, and DRAM.

SSD Offload via LMCache

KV cache disk offloading for NVIDIA Dynamo + vLLM, reducing GPU and DRAM pressure for large models and long contexts.

vLLM & SGLang Upgrades

vLLM upgraded to 0.19.0 with support for Gemma 4, GLM 5.1, and Minimax 2.7. SGLang upgraded to 0.5.10.post1.

Fully Automated AI Performance Benchmarking

Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.

Workloads and Configurations

Configure and test across multiple dimensions simultaneously.

ConcurrencyISL / OSLTensor ParallelPrecisionQuantizationKV Cache DtypeReasoning ParserStreaming (On/Off)Inference Mode (Chat/Completion)LLM Serving (vLLM, SGLang, NVIDIA NIM, Dynamo, TensorRT-LLM + versions)KV Cache OffloadMultinode InferenceGPU Count

Chips

Benchmark across architectures.

NVIDIA DatacenterAMD InstinctIntel Gaudi 3AMD EPYCIntel XeonRTX GPUs

Models

Test the latest foundation models.

Gemma 4MiniMax M3Kimi K2.7NVIDIA Nemotron 3GLM 5.2DeepSeek V4Qwen 3.6GPT OSSLlama 4Mistral

Real-Time Metrics

Track performance across every dimension.

Performance Metrics

Output Tokens

tok/s

Requests

req/s

Words

words/s

TTFT

P50 · P90 · P99

TPOT

P50 · P90 · P99

System Power

Watts

GPU Power

Watts

CPU Power

Watts

Efficiency

tok/s/W

System Metrics

GPU Utilization

%

GPU Memory

Usage

GPU Temperature

°C

CPU Usage

%

CPU Frequency

GHz

Disk I/O

R/W

SSD Metrics

IOPS/Latency

Memory Usage

GB

Real-Time Monitoring

Performance Visualization with Pulse

Real-time visualization of performance metrics across your benchmarking runs.

Pulse Performance Visualization