Cut inference spend40-90%.
Route each LLM request to the cheapest model that still meets your quality bar. Grounded in Metrum benchmark data, not a black-box classifier.
Live Request Routing
Fast & Cheap
Small open model
Balanced
Mid-size model
Frontier
Top reasoning model
Requests Routed
48213
Avg Cost / 1M tok
$0.05
Cost Saved
98%
80.7%
Cost Reduction
$1,200+
Total Savings
3
API Surfaces
10
Routing Controls
How It Works
Routing is a pure, testable function, not a hidden model you have to trust blind.
Prompt analysis
Every request lands on one governed endpoint. The router reads the incoming prompt, tier hint, tool requirements, and image or context payload before anything is sent upstream.
Difficulty classification
A policy, or an external classifier it calls, reads the request and decides how hard it actually is: a short summarization ask and a multi-file refactor do not belong in the same tier.
Model selection
The tier resolves to a ranked pool of candidates across providers, weighted by cost, latency, and measured capability. First eligible target wins; the rest are fallback.
Quality gate
Only models that already cleared that tier's evaluation contract are in the pool. If a candidate fails upstream, the router falls back transparently. No silent quality loss.
Benchmark-grounded routing.
Tier floors come from measured data, not a black box.
Same requests, same task, same success bar - just cheaper.
Provable savings, no quality loss.
A cheaper model that fails silently is not a savings.
Candidate models
Eval gate
Task-specific suite
Tier pool
Cleared only
Re-runs on every model swap. Fail the gate, no production traffic.
Where Smart Router Sits
The router space spans research libraries, managed aggregators, and proxy gateways. Here is where each sits.
| Capability | Metrum Smart Router | LLMRouter Research framework | Not Diamond Learned router | RouteLLM OSS framework | LiteLLM OSS proxy | Azure Model router |
|---|---|---|---|---|---|---|
| Programmable routing policy you own and version | ||||||
| Self-hosted / private model support (vLLM, SGLang) | ||||||
| Cost & token telemetry per user, project, and key | ||||||
| Capacity pooling across multiple provider accounts | ||||||
| Benchmark-validated quality floor per tier | ||||||
| Enterprise self-hosted / air-gapped licensing |
← Swipe to see all comparisons →
Based on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.
One router. Every model.
Drop in the endpoint, write a policy, ship. Pick the path that matches where you are.
Get a demo
Walk through routing policy, quality contracts, and cost reporting with your own workload as the example.
Schedule a demoContact sales
Enterprise self-hosted, private managed, or air-gapped deployment with signed license enforcement.
Talk to salesTransparent per-token pricing, no hidden routing margin. You see exactly what each hop costs.