Every request your systems send to an AI model goes to whichever model somebody configured first. On OpenAI’s standard tier in September 2026, the models alongside that choice are priced at $50, $10, and $0.50 per million output tokens, so identical work can cost a hundredfold more depending only on where it lands.*
Some of what you send genuinely needs the expensive one. A good deal of it does not, and nothing in the configuration separates the two. Finding the candidates takes four things you already know about a workload.
Moving one takes a test. What follows sets out those four, the two cases where they point in opposite directions, and why the tooling built to automate the decision is the part that goes wrong.
How your default model got chosen
The model running your workloads today was picked by whoever built the first workload, under conditions that have since changed. It then spread across everything built afterward, because no line item records which model was selected or why. By the time a renewal conversation starts, the largest variable in the bill has already been fixed, months earlier, by an engineer solving a different problem.
What keeps it in place is an assumption that models get cheaper on their own, so the decision corrects itself in time.
What the research found.
The correction is uneven. Anthropic’s Opus line moved between generations from $15 in and $75 out per million tokens to $4 and $20. Its Haiku line moved the other way, from $0.80 and $4 to $1 and $5.
Prices are falling at the top of the range and rising at the bottom, which is the end you would be moving toward.
Case in point
In 2024, Checkr moved a background-check classification task with 230 possible responses from GPT-4 to a Llama-3 model with around eight billion parameters, fine-tuned on its own historical checks.
Monthly cost fell from roughly $12,000, or about $7,000 with retrieval attached, to around $800. Accuracy rose for both clean and messy records, and response time dropped from seconds to half a second.
Those are Checkr’s own 2024 figures, and the change involved training a model on their data, so the ratio is what travels, and the method is its own undertaking.
The work never changed, only the thing answering it did, which raises the question of how you tell, across your own workloads, where that swap is safe.
Stay updated with Simform’s weekly insights.
Which work moves and which stays
The exercise is a sort. Rank your workloads by request volume, take the top of that list, and ask four things of each. How many requests it handles in a month, how much the work varies between them, what a wrong answer costs to put right, and whether someone is waiting while it runs. Volume tells you where to look first. The other three tell you whether moving is safe.
That yields one rule. High-volume work with repeatable output and cheap, reversible errors is the first candidate to move. High-consequence work stays where it is unless a test shows the quality gap is immaterial. Low-spend work stays because the analysis costs more than the savings.
Where the answer is move.
Take inbound records that need tagging and routing. Volume is high, the output is narrow, and a mistake costs a re-queue. Nothing there requires reasoning across several dependent steps, so the premium buys you nothing.
Peer-reviewed work presented at ICLR 2025 measured more than a twofold cost reduction with no fall in measured response quality, on benchmarks where the cheap and expensive options sat fifty times apart in price.
Move the workload, put human review on a sampled slice that includes your messy and long-tail cases, and watch accuracy alongside cost per thousand requests.
Where the answer is stay.
Now take a drafted response that becomes a customer commitment or enters a regulated record. The volume looks tempting, and the answer is still no, because that workflow has a narrow tolerance for error, and a cheaper model only holds if its mistakes stay inside it.
When a downgrade produces more corrections or more regulatory exposure, token savings stop being the number that matters.
Running this by hand once takes a morning. Running it continuously is where teams reach for a router.
When routing is worth the trouble
Routing earns its overhead when requests genuinely differ in what they need. Where most of them look alike, pick one appropriately sized model and skip the machinery.
The machinery is not free. RouterArena, which benchmarks routers against one another, found current routers inefficient at reaching cheaper models, including one whose pool was restricted to a single vendor’s family.
Microsoft is direct about the same trade, noting that a lower estimated cost does not justify a quality regression and that any request needing the same model every time should keep a direct deployment. Its architecture guidance adds that dynamic routing makes spend harder to forecast.
What this costs you.
Three controls.
A spend cap per team or application, enforced before the bill arrives.
A quality or correction-rate threshold agreed before rollout, so a workload that crosses it gets pinned back to the previous model while someone investigates.
And a named person who hears when the model pool changes, since an automatically updated pool can shift cost and quality with nobody shipping code.
Quality on edge cases tends to degrade first, which nobody sees unless they were asked to watch.
Put the decision into practice
Rank by request volume and take the top five. Test a smaller model against representative traffic, including the messy cases, and hold back anything whose output carries a commitment or a regulatory obligation. Agree on the threshold that sends a workload back. Set the cap, and name the owner.
Build this as a repeatable exercise because the answer keeps moving. Vendor prices shift several times a year, and each shift quietly changes which workloads sit on the wrong side of the line.
Run it on your highest-volume workload this week, and you will know whether the rest of the list is worth opening.
*Prices checked against OpenAI’s and Anthropic’s published pricing pages on 30 September 2026. Both change frequently.