Webinar

Turning Unified Patient Data into Actionable Clinical Intelligence

📅 October 15, 2026 | 9–10 AM PT · 12–1 PM ET

Register Now

Summarize with AI

Not enough time? get the key points instantly.

Every request your systems send to an AI model goes to whichever model somebody configured first. On OpenAI’s standard tier in September 2026, the models alongside that choice are priced at $50, $10, and $0.50 per million output tokens, so identical work can cost a hundredfold more depending only on where it lands.*

Some of what you send genuinely needs the expensive one. A good deal of it does not, and nothing in the configuration separates the two. Finding the candidates takes four things you already know about a workload.

Moving one takes a test. What follows sets out those four, the two cases where they point in opposite directions, and why the tooling built to automate the decision is the part that goes wrong.

How your default model got chosen

The model running your workloads today was picked by whoever built the first workload, under conditions that have since changed. It then spread across everything built afterward, because no line item records which model was selected or why. By the time a renewal conversation starts, the largest variable in the bill has already been fixed, months earlier, by an engineer solving a different problem.

What keeps it in place is an assumption that models get cheaper on their own, so the decision corrects itself in time.

What the research found.

The correction is uneven. Anthropic’s Opus line moved between generations from $15 in and $75 out per million tokens to $4 and $20. Its Haiku line moved the other way, from $0.80 and $4 to $1 and $5.

Prices are falling at the top of the range and rising at the bottom, which is the end you would be moving toward.

Case in point

In 2024, Checkr moved a background-check classification task with 230 possible responses from GPT-4 to a Llama-3 model with around eight billion parameters, fine-tuned on its own historical checks.

Monthly cost fell from roughly $12,000, or about $7,000 with retrieval attached, to around $800. Accuracy rose for both clean and messy records, and response time dropped from seconds to half a second.

Those are Checkr’s own 2024 figures, and the change involved training a model on their data, so the ratio is what travels, and the method is its own undertaking.

The work never changed, only the thing answering it did, which raises the question of how you tell, across your own workloads, where that swap is safe.

Stay updated with Simform’s weekly insights.

Which work moves and which stays

The exercise is a sort. Rank your workloads by request volume, take the top of that list, and ask four things of each. How many requests it handles in a month, how much the work varies between them, what a wrong answer costs to put right, and whether someone is waiting while it runs. Volume tells you where to look first. The other three tell you whether moving is safe.

That yields one rule. High-volume work with repeatable output and cheap, reversible errors is the first candidate to move. High-consequence work stays where it is unless a test shows the quality gap is immaterial. Low-spend work stays because the analysis costs more than the savings.

Where the answer is move.

Take inbound records that need tagging and routing. Volume is high, the output is narrow, and a mistake costs a re-queue. Nothing there requires reasoning across several dependent steps, so the premium buys you nothing.

Peer-reviewed work presented at ICLR 2025 measured more than a twofold cost reduction with no fall in measured response quality, on benchmarks where the cheap and expensive options sat fifty times apart in price.

Move the workload, put human review on a sampled slice that includes your messy and long-tail cases, and watch accuracy alongside cost per thousand requests.

Where the answer is stay.

Now take a drafted response that becomes a customer commitment or enters a regulated record. The volume looks tempting, and the answer is still no, because that workflow has a narrow tolerance for error, and a cheaper model only holds if its mistakes stay inside it.

When a downgrade produces more corrections or more regulatory exposure, token savings stop being the number that matters.

Running this by hand once takes a morning. Running it continuously is where teams reach for a router.

When routing is worth the trouble

Routing earns its overhead when requests genuinely differ in what they need. Where most of them look alike, pick one appropriately sized model and skip the machinery.

The machinery is not free. RouterArena, which benchmarks routers against one another, found current routers inefficient at reaching cheaper models, including one whose pool was restricted to a single vendor’s family.

Microsoft is direct about the same trade, noting that a lower estimated cost does not justify a quality regression and that any request needing the same model every time should keep a direct deployment. Its architecture guidance adds that dynamic routing makes spend harder to forecast.

What this costs you.

Three controls.

A spend cap per team or application, enforced before the bill arrives.

A quality or correction-rate threshold agreed before rollout, so a workload that crosses it gets pinned back to the previous model while someone investigates.

And a named person who hears when the model pool changes, since an automatically updated pool can shift cost and quality with nobody shipping code.

Quality on edge cases tends to degrade first, which nobody sees unless they were asked to watch.

Put the decision into practice

Rank by request volume and take the top five. Test a smaller model against representative traffic, including the messy cases, and hold back anything whose output carries a commitment or a regulatory obligation. Agree on the threshold that sends a workload back. Set the cap, and name the owner.

Build this as a repeatable exercise because the answer keeps moving. Vendor prices shift several times a year, and each shift quietly changes which workloads sit on the wrong side of the line.

Run it on your highest-volume workload this week, and you will know whether the rest of the list is worth opening.

*Prices checked against OpenAI’s and Anthropic’s published pricing pages on 30 September 2026. Both change frequently.

Stay updated with Simform’s weekly insights.

Hiren is CTO at Simform with an extensive experience in helping enterprises and startups streamline their business performance through data-driven innovation.

Sign up for the free Newsletter

For exclusive strategies not found on the blog

Revisit consent button
How we use your personal information

We do not collect any information about users, except for the information contained in cookies. We store cookies on your device, including mobile device, as per your preferences set on our cookie consent manager. Cookies are used to make the website work as intended and to provide a more personalized web experience. By selecting ‘Required cookies only’, you are requesting Simform not to sell or share your personal information. However, you can choose to reject certain types of cookies, which may impact your experience of the website and the personalized experience we are able to offer. We use cookies to analyze the website traffic and differentiate between bots and real humans. We also disclose information about your use of our site with our social media, advertising and analytics partners. Additional details are available in our Privacy Policy.

Required cookies Always Active

These cookies are necessary for the website to function and cannot be turned off.

Optional cookies

Under the California Consumer Privacy Act, you may choose to opt-out of the optional cookies. These optional cookies include analytics cookies, performance and functionality cookies, and targeting cookies.

Analytics cookies

Analytics cookies help us understand the traffic source and user behavior, for example the pages they visit, how long they stay on a specific page, etc.

Performance cookies

Performance cookies collect information about how our website performs, for example,page responsiveness, loading times, and any technical issues encountered so that we can optimize the speed and performance of our website.

Targeting cookies

Targeting cookies enable us to build a profile of your interests and show you personalized ads. If you opt out, we will share your personal information to any third parties.