Skip to main content
Metrum Smart Router
Live

Live Request Routing

Request
Simple summarization

Fast & Cheap

Small open model

$0.05/M

Balanced

Mid-size model

$0.35/M

Frontier

Top reasoning model

$2.10/M

Requests Routed

48213

Avg Cost / 1M tok

$0.05

Cost Saved

98%

80.7%

Cost Reduction

$1,200+

Total Savings

3

API Surfaces

10

Routing Controls

How It Works

Routing is a pure, testable function, not a hidden model you have to trust blind.

01

Prompt analysis

Every request lands on one governed endpoint. The router reads the incoming prompt, tier hint, tool requirements, and image or context payload before anything is sent upstream.

02

Difficulty classification

A policy, or an external classifier it calls, reads the request and decides how hard it actually is: a short summarization ask and a multi-file refactor do not belong in the same tier.

03

Model selection

The tier resolves to a ranked pool of candidates across providers, weighted by cost, latency, and measured capability. First eligible target wins; the rest are fallback.

04

Quality gate

Only models that already cleared that tier's evaluation contract are in the pool. If a candidate fails upstream, the router falls back transparently. No silent quality loss.

Benchmark-grounded routing.

Tier floors come from measured data, not a black box.

Lower cost
Direct to a frontier model
Metrum, routed

Same requests, same task, same success bar - just cheaper.

Provable savings, no quality loss.

A cheaper model that fails silently is not a savings.

Candidate models

Eval gate

Task-specific suite

Tier pool

Cleared only

Re-runs on every model swap. Fail the gate, no production traffic.

Every hop logged
Versioned TypeScript policy

Where Smart Router Sits

The router space spans research libraries, managed aggregators, and proxy gateways. Here is where each sits.

CapabilityMetrum
Smart Router
LLMRouter
Research framework
Not Diamond
Learned router
RouteLLM
OSS framework
LiteLLM
OSS proxy
Azure
Model router
Programmable routing policy you own and version
Self-hosted / private model support (vLLM, SGLang)
Cost & token telemetry per user, project, and key
Capacity pooling across multiple provider accounts
Benchmark-validated quality floor per tier
Enterprise self-hosted / air-gapped licensing

← Swipe to see all comparisons →

Based on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.

One router. Every model.

Drop in the endpoint, write a policy, ship. Pick the path that matches where you are.

Get a demo

Walk through routing policy, quality contracts, and cost reporting with your own workload as the example.

Schedule a demo

Contact sales

Enterprise self-hosted, private managed, or air-gapped deployment with signed license enforcement.

Talk to sales

Transparent per-token pricing, no hidden routing margin. You see exactly what each hop costs.