Skip to main content
Metrum Router
Live

Live Request Routing

Request
Simple summarization

Fast & Cheap

Small open model

$0.05/M

Balanced

Mid-size model

$0.35/M

Frontier

Top reasoning model

$2.10/M

Requests Routed

48213

Avg Cost / 1M tok

$0.05

Cost Saved

98%

Team size1 dev
1
10
25
50
100

Saved per year

~$100.0K

~6.7x cheaper than frontier-only
Frontier-only $117,643/devRouted $17,646/dev
Apache-2.0 open-source
One governed API
Hosted + private
Evidence built in

Frontier intelligence is valuable.
Frontier-everywhere is wasteful.

Most agent workflows are not one call to a premium model - they are dozens: reading context, parsing tool results, orchestrating subagents, retrying on errors. Paying frontier prices for every one of those steps is the default today, not a requirement.

How It Works

Routing is a pure, testable function, not a hidden model you have to trust blind.

01

Receive

Every request lands on one governed endpoint using the API you already speak - OpenAI Chat, OpenAI Responses, or Anthropic Messages.

02

Filter

The router checks what the request actually needs: dialect, tool use, images, reasoning, structured output, payload size, token caps, and access.

03

Select

Your deployment policy - weighted, failover, scored, TypeScript, or an external service - picks the target from the pool of requests that clear the filter.

04

Explain

Every hop is recorded as evidence: cost, latency, throughput, attempts, errors, cache hits, and fallback. Nothing is a black box after the fact.

Old hardware and new hardware, one endpoint.

Route smaller models to older or lower-cost enterprise hardware, and larger models to premium accelerators - all behind a single governed endpoint your applications never have to know about.

Register validated vLLM, SGLang, or other compatible private inference services as routing targets, and the policy decides which requests reach which hardware.

Cost-efficient hardware

Smaller models

Premium accelerators

Larger models

One governed endpoint

Benchmark-grounded routing.

Tier floors come from measured data, not a black box.

Lower cost
Direct to a frontier model
Metrum, routed

Same requests, same task, same success bar - just cheaper.

Provable savings, no quality loss.

A cheaper model that fails silently is not a savings.

Candidate models

Eval gate

Task-specific suite

Tier pool

Cleared only

Promote, hold, split, or roll back groups and targets based on evidence - not model reputation alone.

Every hop logged
Versioned TypeScript policy

Open source by design

Open source. Self-hostable. Built to be inspected.

Metrum Router is Apache-2.0 open-source software. Review the behavior, run it in your environment, extend deployment-owned policy, and validate it against your workloads.

model_group: enterprise-coding

policy: deployment-owned

eligible_targets:

- private-efficient

- hosted-workhorse

- frontier-reasoning

quality_gate: required

evidence:

- request_time_cost

- latency_and_throughput

- attempts_and_fallback

Everything the router owns

Routing and policy are the surface. Underneath, every request is attributed, shaped, protected, and audited.

Routing & Policy

  • Static, weighted, failover, dynamic-score, TypeScript, or external-policy routing
  • Quality and capability contracts enforced before selection
  • Stable model groups across providers, accounts, and endpoints

Cost Attribution

  • Tracked by owner, caller key, public token ID, project, environment, client, and model group
  • Request-time cost and baseline-savings attribution

Traffic Shaping

  • Request and token burst controls
  • RPM, TPM, and concurrency limits
  • Per-caller overrides, bounded queue or fail-fast

Upstream Protection

  • Provider, model, and target shaping
  • Adaptive 429 / quota backoff
  • Capacity pooling, noisy-neighbor and fallback-storm protection

Budget Enforcement

  • Daily, monthly, and lifetime budgets
  • Input plus requested-output reservation
  • In-flight concurrency protection, key lifecycle management

Compatibility Gates

  • Chat, Responses, and Messages support
  • Tools, structured outputs, images / VLM, reasoning
  • Dialect-matched target selection, output-cap enforcement

Operations & Observability

  • Latency, TTFB, throughput, attempts, fallback, cache, errors
  • Request IDs joining usage, traces, shapes, and diagnostics
  • CLI and authenticated browser reports with relational rollups

Administration

  • Server-side provider keys
  • Basic Auth or OIDC identity
  • Scoped authorization, metrics isolation
  • Optional PII filtering and retention controls

Where Metrum Router Sits

The router space spans research libraries, managed aggregators, and proxy gateways. Here is where each sits.

Compare Metrum vs.

Programmable routing policy you own and version

Metrum AIRouterFull support
LLMRouterFull support

Self-hosted / private model support (vLLM, SGLang)

Metrum AIRouterFull support
LLMRouterFull support

Cost & token telemetry per user, project, and key

Metrum AIRouterFull support
LLMRouterNone

Capacity pooling across multiple provider accounts

Metrum AIRouterFull support
LLMRouterFull support

Benchmark-validated quality floor per tier

Metrum AIRouterFull support
LLMRouterPartial

Enterprise self-hosted / air-gapped licensing

Metrum AIRouterFull support
LLMRouterNone

Based on publicly available product documentation as of July 2026. Dash-circle = partial / conditional support.

Give every team one AI endpoint. Keep control behind it.

Drop in the endpoint, write a policy, ship - self-hosted, or with our help. Pick the path that matches where you are.

Get a demo

Walk through routing policy, quality contracts, and cost reporting with your own workload as the example.

Schedule a demo

Contact sales

Enterprise self-hosted, private managed, or air-gapped deployment with signed license enforcement.

Talk to sales

View on GitHub

Apache-2.0 licensed. Read the source, self-host it, or open an issue.

View on GitHub

Explore documentation

API compatibility, routing policy customization, quality evaluation, and more.

Read the docs

Transparent per-token pricing, no hidden routing margin. You see exactly what each hop costs.