One AI endpoint.The best model for every job.
Route every request across hosted and private models by capability, quality, cost, latency, and your policy. Developers keep one stable API while platform teams own the model mix - grounded in Metrum benchmark data, not a black-box classifier.
Works with
Live Request Routing
Fast & Cheap
Small open model
Balanced
Mid-size model
Frontier
Top reasoning model
Requests Routed
48213
Avg Cost / 1M tok
$0.05
Cost Saved
98%
Saved per year
~$100.0K
Frontier intelligence is valuable.
Frontier-everywhere is wasteful.
Most agent workflows are not one call to a premium model - they are dozens: reading context, parsing tool results, orchestrating subagents, retrying on errors. Paying frontier prices for every one of those steps is the default today, not a requirement.
How It Works
Routing is a pure, testable function, not a hidden model you have to trust blind.
Receive
Every request lands on one governed endpoint using the API you already speak - OpenAI Chat, OpenAI Responses, or Anthropic Messages.
Filter
The router checks what the request actually needs: dialect, tool use, images, reasoning, structured output, payload size, token caps, and access.
Select
Your deployment policy - weighted, failover, scored, TypeScript, or an external service - picks the target from the pool of requests that clear the filter.
Explain
Every hop is recorded as evidence: cost, latency, throughput, attempts, errors, cache hits, and fallback. Nothing is a black box after the fact.
Old hardware and new hardware, one endpoint.
Route smaller models to older or lower-cost enterprise hardware, and larger models to premium accelerators - all behind a single governed endpoint your applications never have to know about.
Register validated vLLM, SGLang, or other compatible private inference services as routing targets, and the policy decides which requests reach which hardware.
Cost-efficient hardware
Smaller models
Premium accelerators
Larger models
One governed endpoint
Benchmark-grounded routing.
Tier floors come from measured data, not a black box.
Same requests, same task, same success bar - just cheaper.
Provable savings, no quality loss.
A cheaper model that fails silently is not a savings.
Candidate models
Eval gate
Task-specific suite
Tier pool
Cleared only
Promote, hold, split, or roll back groups and targets based on evidence - not model reputation alone.
Open source by design
Open source. Self-hostable. Built to be inspected.
Metrum Router is Apache-2.0 open-source software. Review the behavior, run it in your environment, extend deployment-owned policy, and validate it against your workloads.
model_group: enterprise-coding
policy: deployment-owned
eligible_targets:
- private-efficient
- hosted-workhorse
- frontier-reasoning
quality_gate: required
evidence:
- request_time_cost
- latency_and_throughput
- attempts_and_fallback
Everything the router owns
Routing and policy are the surface. Underneath, every request is attributed, shaped, protected, and audited.
Routing & Policy
- Static, weighted, failover, dynamic-score, TypeScript, or external-policy routing
- Quality and capability contracts enforced before selection
- Stable model groups across providers, accounts, and endpoints
Cost Attribution
- Tracked by owner, caller key, public token ID, project, environment, client, and model group
- Request-time cost and baseline-savings attribution
Traffic Shaping
- Request and token burst controls
- RPM, TPM, and concurrency limits
- Per-caller overrides, bounded queue or fail-fast
Upstream Protection
- Provider, model, and target shaping
- Adaptive 429 / quota backoff
- Capacity pooling, noisy-neighbor and fallback-storm protection
Budget Enforcement
- Daily, monthly, and lifetime budgets
- Input plus requested-output reservation
- In-flight concurrency protection, key lifecycle management
Compatibility Gates
- Chat, Responses, and Messages support
- Tools, structured outputs, images / VLM, reasoning
- Dialect-matched target selection, output-cap enforcement
Operations & Observability
- Latency, TTFB, throughput, attempts, fallback, cache, errors
- Request IDs joining usage, traces, shapes, and diagnostics
- CLI and authenticated browser reports with relational rollups
Administration
- Server-side provider keys
- Basic Auth or OIDC identity
- Scoped authorization, metrics isolation
- Optional PII filtering and retention controls
Where Metrum Router Sits
The router space spans research libraries, managed aggregators, and proxy gateways. Here is where each sits.
| Capability | LLMRouter Research framework | Not Diamond Learned router | RouteLLM OSS framework | LiteLLM OSS proxy | Azure Model router | |
|---|---|---|---|---|---|---|
| Programmable routing policy you own and version | ||||||
| Self-hosted / private model support (vLLM, SGLang) | ||||||
| Cost & token telemetry per user, project, and key | ||||||
| Capacity pooling across multiple provider accounts | ||||||
| Benchmark-validated quality floor per tier | ||||||
| Enterprise self-hosted / air-gapped licensing |
Compare Metrum vs.
Programmable routing policy you own and version
Self-hosted / private model support (vLLM, SGLang)
Cost & token telemetry per user, project, and key
Capacity pooling across multiple provider accounts
Benchmark-validated quality floor per tier
Enterprise self-hosted / air-gapped licensing
Based on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.
Give every team one AI endpoint. Keep control behind it.
Drop in the endpoint, write a policy, ship - self-hosted, or with our help. Pick the path that matches where you are.
Get a demo
Walk through routing policy, quality contracts, and cost reporting with your own workload as the example.
Schedule a demoContact sales
Enterprise self-hosted, private managed, or air-gapped deployment with signed license enforcement.
Talk to salesExplore documentation
API compatibility, routing policy customization, quality evaluation, and more.
Read the docsTransparent per-token pricing, no hidden routing margin. You see exactly what each hop costs.