For finance & engineering leadership

You are overpaying for AI. We can prove exactly where.

Add one line of code, and get a full report card on the price and performance of your AI application.

  • We build evals automatically from your live traffic, then replay them against cheaper models on your own provider accounts.
  • Nothing is recommended until quality provably holds. We are never in your request path.
  • The change arrives as a pull request with the evidence attached.
savings-model illustrative

$60k / month

Annual run-rate: $720k

Estimated annual savings

$108k to $288k

Used to scope the assessment. No newsletter, no sharing.

The problem

Nobody owns the question “are we running the right model?”

Picking the right model is a laborious, time-consuming process that requires auditing logs, building evals and running tests.

The decision gets made once

A model is chosen the day a feature ships and stays chosen. Prices move every month. Nothing in your process is watching.

“We chose that model 14 months ago. It's probably wrong now. Nobody's checked.”

Known savings sit in a backlog

Most engineers can point straight at the waste. It's two weeks of work, which is exactly why it loses to revenue work every quarter.

“Free money. It still hasn't made the list.”

The downside risk stops everyone

Switching without evidence bets product quality on a hunch. So the expensive model wins by default, and keeps winning.

“Pull one lever, improve five characteristics, degrade fifteen others.”

Quotes from customers spending $10k to $200k+/mo on LLMs

The shape of the deal

A real cut to the bill that barely touches your engineers.

15 to 40%

Modelled spend reduction

What we model once the routes worth optimizing have been optimized.

~1 line of code

To start measuring

Wrap the client your team already uses. Nothing else in your stack changes.

0

Requests routed through us

Never in your request path. If we went down, your product wouldn't notice.

the 15 to 40% figure is our model, not an audited customer result

Where the money is

Model swaps get the attention. Most of the money is in the work nobody schedules.

Each of these is its own project, which is why each one keeps getting deferred.

Lever Typical reduction Why it hasn't happened yet
Model substitution 40 to 70% on the route Needs an eval set nobody maintains, and a quality case nobody can win without data.
Batch API migration ~30% where eligible Two weeks of unglamorous work that loses to the roadmap.
Prompt distillation 50 to 80% on the route Takes a pipeline that teaches a cheap model what the expensive one knows.
Triage routing 30 to 60% blended Safe only once you can measure what the cheap model handles at parity.
Parameter tuning 10 to 30% Reasoning budgets, retries, and fallbacks are invisible without per-feature cost.
Cache strategy 10 to 25% Shows up on no dashboard you own.
Feature cost drivers varies, often large When a provider changes how a feature bills, you find out on the invoice.

Per-route ranges, applied only to the slice of your traffic that qualifies. Not a promise about your total bill.

The spread that makes this worth testing

Below the frontier tier sits a set of models costing a fraction as much per token, and on a lot of routes they hold up fine.

Model Input / 1B Output / 1B Cost index
Frontier tier · where most production traffic sits today
Claude Fable 5 $10,000 $50,000 200
GPT-5.6 Sol $5,000 $30,000 114
Claude Opus 5 $5,000 $25,000 100
Where a lot of that work runs just as well
Kimi K3 $3,000 $15,000 60
Gemini 3.5 Flash $1,500 $9,000 34
GLM-5.2 $760 $2,420 11
DeepSeek V4 Pro $435 $870 5
DeepSeek V4 Flash $90 $180 1

Cheaper per token is not the same as cheaper per task

A model that reasons longer or retries more can cost more per finished answer than the one it replaced. This table is the reason to look, not the answer.

The proof standard

Cheaper is only a saving if quality holds. So we prove it first, in writing.

The objection your CTO will raise, answered

Any vendor can tell you to run a cheaper model. Teams don't, because a quiet drop in quality nobody catches for a month costs more than the savings.

So we replay your real production inputs against the challenger and score every quality characteristic with a confidence interval. A recommendation appears only once that interval clears your parity threshold, and it still has to survive a live canary.

scorecard · ticket-classifier · challenger vs production NON-INFERIOR · −$71,400/yr
−2% margin0+2%
accuracy +0.3
faithfulness +0.1
format validity +0.0
tool-call correctness +0.4
tone −0.6

illustrative scorecard · judge reliability on this route: 0.87 agreement with your labels · bias-probed · sampling stopped at CI convergence

How it works

Four steps. Your team's involvement is step one and a code review.

  1. STEP 01

    Measure

    One line of code attributes every AI call to the feature that made it, in dollars. Usage metadata only.

  2. STEP 02

    Auto-eval

    Eval sets are generated from your own traffic, never hand-written, then replayed against cheaper models inside your allow-lists and residency rules.

  3. STEP 03

    Prove

    A scorecard per feature: quality parity with confidence intervals, plus cost and latency deltas.

  4. STEP 04

    Ship

    A pull request with the evidence attached, rolled out behind a canary, reversible with one flag.

metergraph-overview.mp4 1:23

The risk question

What you are actually exposing by trying this.

FAQ

What leadership asks us first.

Our engineers say they've already optimized this.

Once, on the routes they had budget for. Whether that answer survived six months of price changes is a different question. Run the free layer for a week and you'll know.

What if a cheaper model quietly degrades our product?

Nothing is recommended until the challenger clears non-inferiority testing against your own threshold, judged by a model qualified against your labels. Then it has to survive a live canary, and one flag reverses it.

How much engineering time does this cost us?

An afternoon to wrap your existing client. After that your team reviews pull requests that arrive with their own evidence.

How fast is payback?

Cost per feature shows up within days of instrumenting. Change recommendations follow once there's enough traffic to reach confidence: weeks on high-volume routes, longer on quiet ones.

What does it cost?

The measurement layer is free and open source forever. The hosted engine is a platform fee tiered on traced volume, free for design partners while we calibrate. If the modelled savings don't clear the fee at your volume, we'd rather tell you.

We already have an observability or gateway vendor.

Metergraph sits alongside both. Observability records what your spend was; a gateway routes by rules a human wrote. Neither runs the experiment, and neither opens the pull request.

What about compliance, residency, and sensitive data?

Defaults are metadata-only capture, your own API keys, per-route opt-outs, zero-retention modes, and residency-aware sampling.

For stricter obligations we deploy inside your own VPC, sign a BAA and handle PHI under HIPAA, and provide SOC 2 evidence for your security review.

Where does the 15 to 40% number come from?

It's our model, not an audited customer result: the share of spend that typically sits on optimizable routes, times the per-route ranges above. Your number depends on your traffic mix.

Next step

Find out what your number actually is.

Tell us roughly what you spend and we'll scope an assessment against your real traffic.

savings-assessment

$60k / month

Estimated annual savings

$108k to $288k

Used to scope the assessment. No newsletter, no sharing.

Or start measuring yourself · open source on GitHub ↗ · read the docs