Skip to content
← Back to blog
·7 min read

Stop Paying Frontier Prices for If-Statements

aidevopsmanagement
Stop Paying Frontier Prices for If-Statements

This week I went through the cost ledger of my own AIDLC work. 112 lifecycle instances, about $7,270 of tracked agent spend. Every priced call in roughly 31,000 ledger entries ran on a top-tier model. Opus or Fable. Not one cheaper model anywhere.

The two phases I had tagged mechanical myself, testing and deployment, account for 39% of it. About $2,800 on work I had already labeled as not needing much thought.

The embarrassing part: three days ago I shipped the fix. AIDLC now has a skill_tiers setting that maps each phase tier to its own model and effort level. It's off by default. My own config doesn't set it.

I don't think I'm unusual. Most engineering orgs I talk to do the same thing: one model, the best one, for everything. Picking is a decision. Defaulting isn't. So nobody picks.

Most of your AI calls are if-statements

Look at what your production code actually asks a model. Is this ticket billing or technical? Is this message a jailbreak attempt? Which of these five queues does it go to? Does this document contain personal data? Rate this answer from 1 to 5.

The answer is a label. Sometimes a number. And we pay a model that could write a sonnet about Kubernetes to return the word "billing", then parse the string and hope it spelled it right.

That was defensible when there was no real alternative. There is one now.

Jev and the System One bet

On 15 September 2026, TypeSafe AI put Jev into early access. It's a model that refuses to write. You give it a block of unstructured state plus a set of typed questions (pick one of these options, score this on a scale, is this statement true) and it gives back a probability distribution for each one. No text at all.

TypeSafe calls this class "System One models", after Kahneman's fast, intuitive mode of thinking. Their own one-line description is "unstructured state in, typed probabilistic decisions out."

The numbers they publish:

  • 70 to 500 ms end to end, which they put at 40x to 200x faster than frontier models on what they call System One shaped queries.
  • $0.042 per million input tokens, with output not metered. Their own comparison table puts frontier input at $0.20 to $10 per million, with output on top.

TypeSafe says openly that it can't prove the price isn't subsidized. Fine. Raise it tenfold and it's still a different universe from what I paid to have Opus run deployment checks.

The name is the part I like most. Jev is named after William Stanley Jevons, the economist who noticed in 1865 that more efficient steam engines made Britain burn more coal, not less. That's the Jevons paradox, and it's the honest bet behind the product: make a judgment cheap enough and people will put it everywhere.

This isn't one startup's idea

Nine days later Fastino released GLiNER2.5-Decide, an open-weight decision model of about 340M parameters under Apache 2.0. It runs on a CPU, at around 167 ms per answer on a 48-core Xeon with no GPU. On Fastino's own 17-dataset suite it scored 60.1%, against 57.5% for a Jev configuration Fastino labels JevK5.

So within two weeks we had two decision models, each winning on a test its own maker wrote. That's where this market is right now.

The idea underneath is older than either of them:

  • 2023. Stanford's FrugalGPT sent queries to cheap models first and escalated only when needed. It matched the best single model with up to 98% less cost.
  • 2024. LMSYS released RouteLLM. Its routers cut cost by over 85% on MT Bench while keeping 95% of GPT-4's quality. On GSM8K the saving was only 35%, which tells you how much this depends on the workload.
  • 2025. NVIDIA researchers argued that small language models should be the default for the repetitive, narrow calls agents make all day.

What's new isn't routing. It's that the cheap rung of the ladder is now a different kind of model, not a shrunken copy of the expensive one.

What none of this proves

I'd be doing exactly what I complain about in vendor decks if I stopped there.

Every speed and cost multiple above comes from the people selling the model. TypeSafe's workflow evals were written by its own team, scored against reference answers from GPT-6 Astra and Fable 5.1, and run from laptops on the US West Coast. TypeSafe says most of this itself, which I respect. It's still not independent.

"Can't hallucinate" means "can't return an invalid type". A decision model can still pick the wrong valid answer, confidently. That's an accuracy problem wearing a format guarantee. TrueFoundry's review makes the same point.

Calibration is the whole game, and it's untested. If 0.9 confidence means right about 90% of the time, you can build a router on it. If it doesn't, you've built a very fast way to be wrong. As of that review, nobody outside TypeSafe had checked.

No reasoning, no escape hatch. A classifier handed a ticket about something you never imagined still picks one of your boxes. There's no trace to read afterwards.

And my own numbers deserve the same treatment. Not all of that 39% is if-statements. Deployment in AIDLC writes release notes and decides whether to halt a release, and some of that genuinely needs a strong model. The point isn't that $2,800 was wasted. It's that I never measured which part of it was.

Model choice is an architecture decision

In The Driver Is a Script, Not an Agent I argued that anything with exactly one right answer belongs in code, not in a model. This is the same idea one level up. Inside the work that does need a model, match the model to the shape of the answer.

Three tiers, assigned per call type, not per team:

  1. Reasoning. Open-ended design, ambiguous requirements, anything a human will argue with. Frontier model, high effort.
  2. Standard. Code against a clear spec, summaries, rewrites. A mid-tier model.
  3. Decision. Classify, route, score, extract, guard. A decision model or a small model. Milliseconds, fractions of a cent.

Then the glue: the decision tier returns a confidence, and below a threshold the call moves up a tier or goes to a person. That's FrugalGPT's cascade, except the first rung got two orders of magnitude cheaper.

In AIDLC the whole thing is one block of config (changelog entry):

skill_tiers:
  reasoning: { model: opus, effort: high }
  standard: { model: sonnet, effort: medium }
  mechanical: { model: haiku, effort: low }

Now remember Jevons. When a decision costs almost nothing, you'll put one on every log line, every pull request comment, every inbound email. Unit cost goes down, total spend goes up. That makes getting the tiers right more important over time, not less.

What to do Monday morning

  1. Inventory your model calls by what comes back. Ask each team to list every production model call and mark the output: prose, code, or a label or number. Every label row is a candidate.
  2. Shadow-test the three busiest label calls. Run a decision model (Jev, GLiNER2.5-Decide, or a small model you already host) beside the current one for two weeks. Measure agreement, and read the disagreements by hand.
  3. Check calibration yourself. Bucket answers by confidence and measure accuracy in each bucket. If 0.9 isn't close to 90%, you can't route on it, whatever the launch post says.
  4. Put an owner on the escalation threshold. Below a set confidence, the call goes to a bigger model or a person. Someone's name is on that number.
  5. Report cost per call type, not per vendor. A monthly AI invoice tells you nothing. I only found my 39% because my ledger records which phase every call belonged to (how that works).

The best model is the one that's right for the call. For most calls, that isn't the best model.