Every evaluation of a clinical AI model measures how well it predicts. Someone decided what it should predict much earlier, during implementation. By the time an accuracy figure reaches you, that decision already sits inside the number.
Most clinical conditions carry more than one official definition. Each one exists in code, so a system can apply it without a clinician. Your vendor may have trained the model on one definition and tested it on another. A third may govern what your quality reporting counts. The model cannot tell them apart, and neither can the figure you were shown.
To evaluate a healthcare AI vendor’s accuracy claim, ask the vendor to name the computable definition behind the model’s training data, and the one behind its evaluation. Then ask them to score your last twelve months against it. Compare the patient count they return with what your clinical teams treated for that condition.
Stay updated with Simform’s weekly insights.
The same sepsis model found 60% fewer patients under a different definition
Mass General Brigham ran a single sepsis model across 198,494 encounters at nine hospitals, most of them community and critical access sites. The model never changed. The only thing that changed was which definition of sepsis it was scored against.
Under the definition clinicians use, the model was working on 5,832 patients. Under the version CMS built for quality reporting, 2,366. Same six months, same encounters, same model. The population shrank by roughly 60% without anyone touching the model.
That count determines what the deployment costs you. If the model flags 5,832 patients a year, someone has to review those alerts, and the staffing plan, the alert thresholds, and the return you promised the board were all calculated against that number.
Build the plan on 5,832, discover the model is working on 2,366, and the review workload, the clinical time, and the case for buying it all move at once. The two groups also included different people, since one in seven patients died in the larger group versus one in five in the smaller.
The narrower definition selected a sicker population. It did not select a subset of the larger group.
The accuracy figure changes too, and in the direction you would least expect. The model scored best against the quality-reporting definition, yet under that same definition only about one alert in fifteen turned out to be a real sepsis case.
The other fourteen were false alarms a clinician had to open and dismiss. A vendor quoting the highest accuracy number is quoting a real result, and it happens to describe the configuration that puts the heaviest false-alarm load on your staff.
Can you trust a vendor’s published accuracy number?
Not without seeing how it behaves somewhere like your hospital. Researchers validated Epic’s sepsis model v2 across four US health systems, covering 227,091 encounters. The share of alerts that proved correct ran from one in eight to one in four across sites.
Same model, same version, four different answers. The researchers advised anyone deploying it to run a local validation first, which is another way of saying the published figure describes their hospitals and not yours. Ask for the range and the sites behind any single number you are shown.
Changing the definition after go-live costs more than changing a report filter
Most teams discover the price of this decision only when they try to reverse it. If the definition lives in a reporting filter or a cohort query, moving it takes an afternoon.
If the definition sits inside the model’s training target, the work multiplies. Your team has to relabel history, retrain, reselect thresholds, and revalidate before anything reaches a clinician again. Few organizations know which of the two they have until somebody asks the question.
The cost compounds once the model is live.
Your team tuned the alert thresholds against the count that definition produces, and staffing assumed that volume. If the model feeds quality reporting, the numbers you submit shift when the definition shifts.
Someone then has to explain to a board why a metric moved for reasons that have nothing to do with patient care.
Plenty of organizations still have this decision ahead of them. Federal data puts predictive AI use at 37% among independent hospitals against 86% among system members, so for a large part of the market the first clinical model has yet to arrive.
Getting the definition question right before it does is considerably cheaper than unpicking it afterward.
What should a health system ask a healthcare AI vendor before signing?
- Which computable definition was the model trained on, by name and version?
- Was it evaluated against that same definition, or a different one?
- Will you score our last twelve months against it and report the patient count back to us?
- Which definition does your next product use for this same condition?
Build the business case on the count from the third question. It should reconcile against how many patients your clinical teams treated for that condition, and a gap there costs far less to find during evaluation than in a dashboard six months in.
If the fourth answer names a different definition, you are looking at a governance problem that will outlast this purchase.
One definition per condition, owned and referenced by every system that needs it
Auditing a single model tells you what happened once. It does nothing about the next model, the next vendor, or the next quality measure arriving with its own definition and its own patient count.
Until recently, whichever definition you picked affected only your own dashboards. That is changing. CDC and CMS have built an adult community-onset sepsis mortality measure that scores hospitals against a national benchmark using the CDC surveillance definition, which may well differ from the one your model or your quality team adopted.
A related measure of hospital sepsis programs was submitted to the 2026 CMS Measures Under Consideration list. Once a measure like that is in force, your sepsis performance becomes a public figure calculated on somebody else’s definition, and any gap against your internal number is yours to explain.
Who owns clinical definitions across models and quality measures?
One authoritative definition per condition ends the recurring argument. Write down its logic, purpose, owner, version, and validation evidence. Every system that needs the definition then points at that one record.
Ownership belongs with whoever can say no to a vendor implementation, which in most organizations means data governance and clinical informatics working together, since neither can hold the line alone.
A common data model makes that definition portable. Map clinical data into OMOP against standardized vocabularies, and your team writes the cohort logic once. That logic then runs across every system without a rewrite. Delta-format lineage records when a definition changed and what moved with it.
The same discipline sits underneath patient identity resolution across EHR and claims, where an unowned decision surfaces the same way and just as late. Working with Simform on a Fabric data foundation, a global CDMO cut manual reporting effort by half once definitions stopped being rebuilt for every report. However, that was a life sciences reporting estate, not a provider clinical one. The registry that sits on top of the platform is yours to build, and it is the part no product ships.
Accuracy is the wrong first question for your next evaluation. Ask which definition accuracy was measured against, then ask who in your organization owns that definition for every other system that will need it; if the answer is nobody, start there before the next model arrives.
