Chapter 5.3
AI Health Economics
Artificial intelligence changes what a health technology is — a product that learns, degrades, and can quietly redistribute care — so evaluating it well means pricing not only its accuracy but its upkeep, its effect on the workforce, and the equity it gives or takes away.
Why this matters in health economics
Artificial intelligence (AI) has moved from demonstration to deployment in health systems worldwide, and the money now follows it: payers, ministries, and provider groups are being asked to buy diagnostic algorithms, triage tools, and decision-support systems on the promise that they cut cost or improve outcomes. The promise is often real, but the evidence supporting it is frequently thinner than for a drug or a device, and the economics behave differently. A busy director who applies the reflexes of a standard purchase — fixed price, fixed performance, one-off appraisal — will misprice these technologies in ways that surface only later, as maintenance bills, workforce disputes, or an equity scandal.
The stakes are the ordinary stakes of health economics, sharpened. Public money is spent on tools whose benefit depends on how clinicians actually use them, not on how they perform in a published study. Safety depends on performance that can silently decay as populations and practice change. And equity is exposed in a specific new way: an algorithm trained on one population can systematically underserve another, turning a procurement decision into a distributional one. As of the mid-2020s, regulators such as the United States Food and Drug Administration (FDA) and the European Union have begun building frameworks for these products, but reimbursement and evaluation practice lag behind the technology, leaving decision-makers to improvise.
This chapter is about that gap. It does not re-teach economic evaluation — the cost-effectiveness machinery lives in Chapter 2.1 — Economic Evaluation, and the modelling of uncertainty in Chapter 2.2 — Modelling — nor does it re-cover general digital tools such as apps and telehealth, which belong to Chapter 5.2 — Digital Health Economics. It owns the economics that are specific to AI: evaluating tools that learn, valuing the substitution or augmentation of clinical labour, and treating algorithmic bias as a cost that falls on real people (the equity concepts themselves are Chapter 3.4 — Equity).
Core concepts
Artificial intelligence in health care refers to the use of computational systems — most consequentially those built with machine learning, where models are trained on data rather than explicitly programmed — to perform tasks that would otherwise require clinical judgement. The two families that matter most economically are AI diagnostics (an algorithm that reads a scan, a slide, or a signal and returns a finding) and the clinical decision support system (a tool that recommends, ranks, or flags to a human decision-maker). Both are usually regulated as medical software — software intended for a medical purpose — which the FDA and others term "software as a medical device".
The first concept that reshapes the economics is that an AI product is often not fixed. A conventional device performs the same on the day you decommission it as on the day you installed it; a learning model can change, either because the vendor retrains it or because the world around it shifts. Concept drift — and the related problem of dataset shift, where the population the model meets in service diverges from the population it was trained on — means performance can decay silently. A technology that must be monitored, revalidated, and periodically retrained carries a stream of ongoing costs and risks that a one-off appraisal will miss.
The second concept is the labour economics of automation. AI can substitute for clinical labour (replacing a task a human did) or augment it (making a human faster or better). The distinction is economic, not rhetorical: substitution promises workforce savings but concentrates them, redistributes work, and can deskill a profession; augmentation promises quality or throughput gains but often adds cost rather than removing it. Which one you are actually buying determines where the value comes from — and whether the business case rests on cashable savings that may never materialize if the freed time is simply reabsorbed.
The third concept is automation bias: the human tendency to over-trust a machine's output, accepting a wrong recommendation because it came from a computer. This matters for evaluation because the real-world effectiveness of a decision-support tool depends on the human–machine team, not the algorithm alone. A tool that is accurate in isolation but induces automation bias, or is ignored through alert fatigue, will not deliver its modelled benefit.
The fourth concept is algorithmic bias as an equity cost. When a model performs worse for a subgroup — because that group was under-represented in training data, or because a biased proxy was used as a label — it can systematically deny that group the benefit others receive, or expose them to harm. In economic terms this is a distributional loss that a headline average-effectiveness figure hides entirely; a tool can be cost-effective on average while making the worst-off worse off. Naming and measuring that loss is core to evaluating AI honestly (the measurement machinery for equity is Chapter 3.4 — Equity).
The evaluation frame remains health technology assessment (HTA) — the systematic appraisal of a technology's costs, benefits, and broader consequences — but AI stretches it. Bodies such as England's National Institute for Health and Care Excellence (NICE), Germany's IQWiG, and Canada's CDA-AMC are adapting their methods to technologies that update, and the EU Artificial Intelligence Act (adopted 2024) classifies many medical AI systems as high-risk, adding a layer of compliance cost the business case must carry.
Best practices
Decide first whether you are buying substitution or augmentation, because the value comes from different places. A substitution case rests on cashable savings — fewer reader-hours, closed sessions, redeployed staff — which only count if the freed capacity is actually removed or reused, not quietly reabsorbed. An augmentation case rests on better outcomes or higher throughput, usually at added cost. State which one you are claiming, and hold the business case to the evidence that supports that specific claim rather than a blend of both.
Cost the whole life-cycle, not the licence. The purchase price of a medical algorithm is often the smallest line. Add integration with clinical systems, clinician training, the human oversight the workflow now requires, ongoing performance monitoring, periodic revalidation and retraining, incident investigation, and the compliance burden of regimes such as the EU AI Act. A tool that looks cheap against a radiologist's salary can look expensive once its true total cost of ownership is on the page.
Evaluate the human–machine team, not the algorithm. A model's standalone sensitivity and specificity are inputs, not the answer. What determines real value is how clinicians act on its outputs — whether they over-trust it (automation bias), ignore it (alert fatigue), or use it well. Insist on evidence from the intended workflow and workforce, and treat impressive in-silico accuracy that has never been tested in service as an unproven claim.
Treat algorithmic bias as a measurable cost, not a compliance footnote. Require performance to be reported broken down by the groups your population actually contains — age, sex, ethnicity, skin tone, comorbidity, language, socio-economic status — not just in aggregate. A tool that raises average outcomes while widening a gap between groups has an equity cost that belongs in the appraisal; make it visible and weigh it, because an average that hides a harmed subgroup is not good enough for a public service (see Chapter 3.4 — Equity).
Price the upkeep of a learning technology, and contract for it. Because performance can drift, budget for continuous monitoring and revalidation as a running cost from day one, and write the maintenance obligations into the contract. Establish before purchase who is accountable when the model degrades, how often it will be revalidated, and what triggers withdrawal — so that "who watches the model" is answered by design rather than discovered after an incident.
Insist on a change-management plan for models that update. A technology that changes after procurement breaks the assumption that appraisal is a one-off event. Prefer arrangements where the scope of future changes is pre-specified and governed — the approach behind the FDA's "predetermined change control plan" for adaptive software — so you know what the vendor may alter without re-approval and what would require a fresh evaluation. Buying an unbounded right to change the model is buying an unpriced risk.
Demand a real comparator and beware the accuracy-to-outcome leap. The relevant question is rarely "is the AI accurate?" but "does it beat current practice on outcomes and cost?" — and higher diagnostic accuracy does not automatically convert into better health or lower spend. More findings can mean more follow-up, more overdiagnosis, and more cost with no benefit. Require evidence that links the tool to downstream outcomes and net resource use against the actual standard of care, not against a strawman (the causal-inference standards are Chapter 2.3 — Health Econometrics).
Match the appraisal's rigour to the decision's stakes and reversibility. A low-risk augmentation tool that a clinician can override cheaply warrants a lighter evaluation than an autonomous diagnostic that determines who gets recalled. Scale the evidence you demand — and the monitoring you commit to — to the consequences of the tool being wrong and the difficulty of undoing it, rather than applying one template to everything (this is priority-setting discipline; see Chapter 3.3 — Rationing).
Model the counterfactual workforce honestly. If the case depends on staff savings, be explicit about what those staff would otherwise do and whether the saving is real. In a system with a workforce shortage, "freed" clinician time is usually reabsorbed by unmet demand, so the benefit is more care delivered, not money saved — a genuine gain, but a different one that changes who in your organization should be signing off the business case.
Localize the evidence before you trust it. A model validated in one country, payer, or patient mix may perform differently in yours because the population, the disease prevalence, the imaging equipment, or the coding practices differ. Require validation on data representative of your own population before deployment, and treat a vendor's headline figures from elsewhere as a hypothesis to test locally, not a result to bank.
Plan the exit at the point of entry. Learning tools fail, drift, lose regulatory standing, or are superseded, and clinical services must keep running when they do. Before deployment, confirm that staff retain the skills and capacity to work without the tool, that data and workflows are not so entangled with one vendor that leaving is prohibitive, and that a withdrawal will not itself harm patients — deskilling is a slow cost that only appears when you try to stop.
Keep AI evaluation inside your existing HTA and governance, extended — not a separate fast lane. The pressure to treat AI as special can become an excuse to bypass the scrutiny applied to any other health technology. Route these decisions through the same economic evaluation, safety, and equity governance you use elsewhere, adding the AI-specific questions of drift, bias, and human-factors rather than replacing the discipline with vendor enthusiasm.
Questions to discuss with your team
Are we buying this tool to replace clinical work or to improve it — and does our business case honestly match that answer? This question forces a choice most business cases blur, claiming workforce savings and quality gains at once while proving neither. If the claim is substitution, press on whether the saving is cashable: will posts actually close, or will freed time be swallowed by the waiting list, making the true benefit "more care" rather than "less cost"? If the claim is augmentation, press on whether the added quality justifies the added cost, because augmentation almost always adds spend. The honest answer names one primary source of value, states who in the organization is accountable for realizing it, and admits which savings are speculative. A team that cannot say plainly what it is buying is not ready to buy.
Who could this algorithm underserve, and would we notice before it caused harm? Algorithmic bias is rarely visible in the headline result; it hides inside an average that looks good. Ask which groups in your population were thin or absent in the training data, whether the label the model learned is a fair proxy or a biased one, and whether you have the data broken down finely enough to detect a gap by ethnicity, sex, skin tone, language, or deprivation. The tension is that collecting and monitoring subgroup performance costs money and effort that a pressured deployment will want to skip. An honest answer specifies the subgroups you will monitor, the threshold at which a gap triggers action, and a candid admission of the groups you currently cannot see — because an equity harm you cannot measure is one you will discover through complaints.
What is our plan for the day this model degrades, changes, or must be switched off? Unlike a conventional device, a learning technology can drift silently, be retrained by its vendor, lose regulatory standing, or simply fail — and the service must continue. Discuss who is accountable for watching performance in service, how often the model is revalidated, and what observable trigger would cause you to withdraw it. Probe whether your clinicians would still be competent to work without it after two years of reliance, and whether your data and workflows are so entangled with one supplier that leaving is unaffordable. The honest answer is a named owner, a monitoring cadence, a withdrawal trigger, and a preserved fallback capability — not a hope that the vendor will keep things working.
Have we costed the whole life-cycle of this tool, or only its purchase price? The licence is usually the smallest line in a medical-AI budget, yet it is the number that dominates the business case that reaches the board. Press the team to name every downstream cost the tool creates: integration with clinical systems, staff training, the standing human oversight the workflow now demands, continuous performance monitoring, periodic revalidation and retraining, incident investigation, and the compliance burden of regimes such as the EU AI Act. The tension is that many of these costs are recurring and diffuse, so they fall on operational budgets that were never consulted, while the saving is booked centrally and up front. An honest answer puts total cost of ownership over the realistic life of the tool against the benefit, names which budgets carry each recurring line, and resists the vendor's instinct to compare a licence fee against a clinician's salary as though nothing else changed.
Are we buying the algorithm, or the way our clinicians will actually use it? A model's standalone sensitivity and specificity describe the tool on a bench, not the human–machine team that will deliver care, and the gap between the two is where modelled benefits quietly disappear. Discuss whether your clinicians are likely to over-trust the output (automation bias), ignore it amid a flood of alerts (alert fatigue), or fold it sensibly into judgement — and what in your workflow, training, and culture pushes toward each. The uncomfortable point is that the same algorithm can help in one service and harm in another purely because of how people around it behave, so evidence from someone else's deployment is a hypothesis about yours, not a result. An honest answer insists on evidence drawn from the intended workflow and workforce, plans the human-factors and change-management work as part of the purchase, and treats impressive in-silico accuracy that has never been tested in service as an unproven claim.
Does this tool actually beat current practice on outcomes and cost, against a real comparator? The seductive question is "is the AI accurate?"; the decision-relevant question is "does it do better than what we do now, on the things that matter, once everything is counted?" Probe whether the comparison is against your genuine standard of care or a weakened strawman, and whether higher diagnostic accuracy has been shown to translate into better health rather than simply more findings, more follow-up, more overdiagnosis, and more cost. The tension is that accuracy is easy to demonstrate and cheap to market, while downstream outcome and net-resource evidence is slow, expensive, and often absent — so the pressure is to accept the former as a proxy for the latter. An honest answer names the real comparator, demands evidence linking the tool to downstream outcomes and net resource use, and is willing to say "not proven yet" when the chain from accuracy to benefit has not been shown (the causal-inference standards are Chapter 2.3 — Health Econometrics).
In practice: a health economics example
The Prairie Health Authority, a fictional single-payer regional health service in western Canada, is considering an AI system to screen for diabetic retinopathy — a leading cause of preventable blindness in people with diabetes. Today, retinal photographs from community clinics are read by a small pool of ophthalmologists and trained graders, and the region cannot keep pace: rural and Indigenous communities in particular face long waits, and some patients are lost to follow-up before their images are ever read. A vendor offers an algorithm that reads the images automatically, referring only positive or ungradable cases to a human. The headline pitch is a workforce saving and faster results.
The authority's health economists start by pinning down what is being bought. The vendor frames it as substitution — the algorithm replaces grader time — but in a system with a screening backlog, the freed grader capacity will be absorbed by the untested backlog, not cashed out. The true benefit, they conclude, is augmentation of capacity: more people screened, sooner, especially in underserved communities, with a plausible reduction in preventable blindness downstream. That reframing changes the whole case. The saving is not a line in the salary budget; it is health gain and equity gain, and it must be valued as such, using the cost-per-quality-adjusted-life-year machinery from Chapter 2.1 — Economic Evaluation rather than a simple staffing subtraction.
They then interrogate the equity dimension directly. The model was trained largely on retinal images from urban populations in another country; its published accuracy is strong in aggregate but has not been tested on the authority's Indigenous and rural patients, whose diabetes prevalence is high and whose imaging is often done on older cameras in variable light. The team refuses to accept the vendor's headline figures as transferable. They commission a local validation on a representative sample and require performance to be reported by community, camera type, and image quality — because a tool that performs worse exactly where the need is greatest would convert a screening programme meant to reduce inequality into one that widens it, an algorithmic-bias cost that would dwarf the licence fee in both harm and reputation.
The whole-life costing surprises the finance director. Beyond the licence, the case must carry integration with the imaging systems, training for clinic staff, a standing arrangement for human over-reading of ungradable images, and — critically — continuous monitoring for concept drift as cameras, populations, and diabetes treatment change over the years. The team writes the maintenance and revalidation obligations into the contract, specifies a predetermined scope for model updates, and names an accountable clinical owner for in-service performance. They also insist on a fallback: enough human grading capacity is retained that a withdrawal of the tool would not blind the service.
The authority approves a phased deployment: local validation first, then a monitored rollout beginning in the highest-need communities, with subgroup performance dashboards reviewed quarterly and a pre-agreed accuracy floor that triggers investigation or withdrawal. The economics are judged favourable — the tool extends scarce specialist capacity to people who were previously missed — but the appraisal makes clear that the value depends on the monitoring and equity safeguards, not the algorithm alone. The lesson the board takes is that an AI purchase is a continuing commitment, not a one-off buy, and that the equity question was not a footnote but the centre of the case.
Four sector lenses
Startup
A health-AI start-up typically has strong evidence of algorithmic accuracy and weak evidence of economic value or real-world effect, and its survival depends on a reimbursement pathway that may not yet exist. The temptation is to sell accuracy as if it were outcome, and to under-price the upkeep, monitoring, and equity work that buyers will eventually demand. Founders who design for evaluability from the start — pre-specifying subgroup performance, building monitoring in, and validating on representative populations — build something a payer can actually buy. The credible pitch is not the best area-under-the-curve on a benchmark; it is a defensible claim about outcomes, cost, and who benefits, tested in a real workflow.
Small business
A small but established provider — a radiology practice, a single-site clinic, a specialist laboratory, or a niche software supplier — buys AI as a settled business decision rather than a bet on survival, and its constraint is the reverse of a start-up's: limited capacity to evaluate, integrate, and monitor a learning tool, not limited product. Such a buyer rarely has a data science team to run a local validation or a governance function to watch for drift, so it is unusually exposed to a vendor's headline figures and to the recurring costs — oversight, revalidation, compliance — that outlast the sale. The sensible posture is to lean on shared or external assurance where it exists — a professional body's evaluation, a national HTA verdict, a group-purchasing arrangement — rather than attempting scrutiny it cannot resource alone. Its distinctive risk is quiet lock-in and deskilling: a small practice that reshapes its workflow around one tool can find that leaving, or working without it, is no longer affordable, so preserving a fallback and reading the contract's exit terms matter more here than the licence price.
Enterprise
A large hospital group or insurer buys AI as a portfolio and lives with the operational reality that vendors gloss over: integration, clinician adoption, alert fatigue, monitoring at scale, and the deskilling risk of relying on tools across many sites. At this scale the substitution-versus-augmentation question is a workforce-strategy question with industrial-relations consequences, and a botched "labour-saving" deployment can cost more in disputes and rework than it ever saved. Enterprises are also where algorithmic bias becomes a systemic liability rather than a single case, so the mature ones build standing governance — model registries, performance dashboards, revalidation schedules — and treat AI oversight as a permanent function, not a project.
Government
A ministry or national payer sets the rules everyone else plays by, through regulation, HTA methods, and reimbursement design, and it carries the failures the market will not price — bias that falls on protected groups, drift that harms safety, and the workforce consequences of automation across a whole system. Its levers are the widest: classifying medical AI by risk (as the EU AI Act does), requiring representative validation and post-market monitoring, and deciding whether and how learning technologies are paid for. Its discipline must be widest too, because a fast lane that waves AI through on vendor accuracy claims, or a reimbursement rule that pays for volume of algorithmic reads, will propagate a mistake across every provider at once. Government's distinctive job is to make evaluation, equity, and monitoring the default conditions of entry rather than optional extras (the policy machinery is Chapter 3.2 — Health Policy).
Common failure modes
Selling accuracy as outcome. Treating a high diagnostic-accuracy figure as proof of health or cost benefit. Fix: require evidence linking the tool to downstream outcomes and net resource use against real current practice.
Booking savings that never cash out. Claiming workforce savings that are quietly reabsorbed by unmet demand. Fix: state whether posts actually close; if freed time is reused, value the case as more care delivered, not money saved.
Ignoring the human in the loop. Evaluating the algorithm alone and missing automation bias and alert fatigue. Fix: evaluate the human–machine team in the intended workflow and workforce.
Averaging away an equity harm. Reporting only aggregate performance and missing a subgroup the tool underserves. Fix: require and monitor performance broken down by the groups your population contains, with an action threshold.
One-off appraisal of a learning tool. Treating an adaptive model like a fixed device and never revalidating it. Fix: budget and contract for continuous monitoring, revalidation, and a change-control plan.
Lock-in and deskilling by neglect. Becoming so dependent on one tool and vendor that withdrawal is impossible and staff can no longer work without it. Fix: preserve fallback capability and plan the exit at the point of entry.
A special fast lane for AI. Bypassing normal HTA, safety, and equity scrutiny because the technology is exciting. Fix: route AI through existing governance, extended with drift, bias, and human-factors questions.
Maturity model
| Dimension | Initiate | Develop | Standardize | Manage | Orchestrate |
|---|---|---|---|---|---|
| Value case | Bought on algorithmic accuracy and vendor claims | Substitution vs augmentation named; some cost data gathered | Whole-life cost and a real comparator appraised through HTA as standard | Value tracked in service against outcomes and net cost, with variances acted on | Value continuously re-tested and rebalanced across the portfolio, with tools disinvested when they stop paying |
| Equity | Only aggregate performance seen; bias unexamined | Subgroup risk acknowledged; ad hoc checks | Representative validation and subgroup reporting required before deployment | Subgroup performance monitored live with action thresholds and published | Equity outcomes managed across tools and communities, with results fed back into procurement and design |
| Learning-tool upkeep | Appraised once and forgotten; drift unmonitored | Monitoring discussed but not resourced | Monitoring, revalidation, and change control budgeted and contracted | Automated drift detection, pre-specified updates, and withdrawal triggers in force | Upkeep orchestrated across the estate and vendors, with shared monitoring and coordinated revalidation cycles |
| Workforce | Labour savings assumed; staff not consulted | Substitution effects mapped; training planned | Counterfactual workforce modelled; deskilling and fallback addressed | Augmentation designed with staff; capacity gains and skills actively managed | Workforce strategy and AI planned together across services, with skills and roles evolved deliberately over time |
| Governance | No AI-specific oversight | Case-by-case approvals | AI folded into standing HTA, safety, and equity governance | Model registry and permanent oversight function with continuous audit | Governance extended across partners and supply chain, shaping standards and accountability beyond the organization |
Checklist
- State explicitly whether the tool substitutes for or augments clinical labour, and match the business case to that claim.
- Cost the full life-cycle — integration, training, oversight, monitoring, revalidation, retraining, and compliance — not just the licence.
- Require evidence from the intended human–machine workflow, not standalone algorithmic accuracy.
- Require performance reported by the subgroups your population contains, with an action threshold for any equity gap.
- Validate the model on data representative of your own population before deployment.
- Confirm a real comparator and evidence that accuracy translates into outcomes and net resource use.
- Budget and contract for continuous monitoring, revalidation, and a predetermined change-control plan.
- Name an accountable owner for in-service performance and a trigger for investigation or withdrawal.
- Model the counterfactual workforce honestly, including whether savings are cashable or reabsorbed.
- Preserve fallback capability and plan the exit, guarding against deskilling and vendor lock-in.
- Route the decision through existing HTA, safety, and equity governance, extended with AI-specific questions.
Key sources
- WHO, Ethics and Governance of Artificial Intelligence for Health — World Health Organization guidance on AI in health, including equity and governance principles.
- U.S. Food and Drug Administration — guidance on Software as a Medical Device (SaMD) and predetermined change control plans for adaptive AI/ML-enabled devices.
- European Union — Artificial Intelligence Act (2024), classifying many medical AI systems as high-risk.
- NICE — Evidence Standards Framework for Digital Health Technologies, and evolving methods for evaluating AI and data-driven technologies (as of the mid-2020s).
- OECD — Health at a Glance and OECD work on AI in health systems, for cross-country context.
- ISPOR — good-practice guidance on the economic evaluation of health technologies, applied to AI-enabled care.
References
- Artificial intelligence in healthcare — Wikipedia — https://en.wikipedia.org/wiki/Artificial_intelligence_in_healthcare
- Machine learning — Wikipedia — https://en.wikipedia.org/wiki/Machine_learning
- Clinical decision support system — Wikipedia — https://en.wikipedia.org/wiki/Clinical_decision_support_system
- Medical software — Wikipedia — https://en.wikipedia.org/wiki/Medical_software
- Concept drift — Wikipedia — https://en.wikipedia.org/wiki/Concept_drift
- Automation — Wikipedia — https://en.wikipedia.org/wiki/Automation
- Deskilling — Wikipedia — https://en.wikipedia.org/wiki/Deskilling
- Automation bias — Wikipedia — https://en.wikipedia.org/wiki/Automation_bias
- Algorithmic bias — Wikipedia — https://en.wikipedia.org/wiki/Algorithmic_bias
- Health technology assessment — Wikipedia — https://en.wikipedia.org/wiki/Health_technology_assessment
- Artificial Intelligence Act — Wikipedia — https://en.wikipedia.org/wiki/Artificial_Intelligence_Act
- Diabetic retinopathy — Wikipedia — https://en.wikipedia.org/wiki/Diabetic_retinopathy
- World Health Organization — Ethics and Governance of Artificial Intelligence for Health — https://www.who.int/publications/i/item/9789240029200
- U.S. Food and Drug Administration — Artificial Intelligence and Machine Learning in Software as a Medical Device — https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-software-medical-device
- NICE — Evidence Standards Framework for Digital Health Technologies — https://www.nice.org.uk/corporate/ecd7