Accelerate YourAI PerformanceTesting
Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.
Metrum Insights 4.0
Unified benchmarking, audit-ready results, deploy anywhere.
Latest Features Available
VectorDB benchmarking, KV cache SSD offloading, and upgraded runtimes with support for the latest models.
VectorDB Bench (Milvus)
End-to-end vector database benchmarking with HNSW and DISKANN indexing, plus new storage metrics including IOPS, latency, throughput, queue depth, and DRAM.
SSD Offload via LMCache
KV cache disk offloading for NVIDIA Dynamo + vLLM, reducing GPU and DRAM pressure for large models and long contexts.
vLLM & SGLang Upgrades
vLLM upgraded to 0.19.0 with support for Gemma 4, GLM 5.1, and Minimax 2.7. SGLang upgraded to 0.5.10.post1.
Fully Automated AI Performance Benchmarking
Configure unlimited combinations of models, software, hardware, and hyperparameters in seconds.
Workloads and Configurations
Configure and test across multiple dimensions simultaneously.
Chips
Benchmark across architectures.
Models
Test the latest foundation models.
Real-Time Metrics
Track performance across every dimension.
Performance Metrics
Output Tokens
tok/s
Requests
req/s
Words
words/s
TTFT
P50 · P90 · P99
TPOT
P50 · P90 · P99
System Power
Watts
GPU Power
Watts
CPU Power
Watts
Efficiency
tok/s/W
System Metrics
GPU Utilization
%
GPU Memory
Usage
GPU Temperature
°C
CPU Usage
%
CPU Frequency
GHz
Disk I/O
R/W
SSD Metrics
IOPS/Latency
Memory Usage
GB
Real-Time Monitoring
Performance Visualization with Pulse
Real-time visualization of performance metrics across your benchmarking runs.

Ready to Accelerate Your AI Performance?
Join industry leaders using Metrum Insights to optimize their AI infrastructure.