Why this matters in health economics

Every avoidable harm has a price, and the health system pays it twice: once in the suffering of the patient, and again in the extra bed-days, repeat procedures, litigation, and lost trust that follow. When a surgical patient acquires a bloodstream infection, the system does not save money — it spends far more than the infection would have cost to prevent. Quality and safety economics is the discipline of making these costs visible, so that leaders stop treating safety as a cost centre competing with productivity and start treating it as a source of released capacity.

This matters because the resources wasted on poor quality are enormous and largely hidden. Analysts distinguish the cost of poor quality — rework, waste, over-treatment, and failure — from the cost of getting care right first time, and in most systems the former is a substantial fraction of total spend. The World Health Organization treats unsafe care as a leading source of avoidable death and disability worldwide, on a scale comparable to major disease burdens, and much of it is preventable with known, low-cost practices.

The stakes are also moral and political. Public money buys care on the understanding that the care will not harm; a healthcare-associated infection or a wrong-site operation is a breach of that trust as well as a budget line. Because quality is hard to observe and easy to game, the way a payer chooses to measure and reward it — through pay-for-performance, value-based purchasing, or penalties for never events — shapes clinical behaviour for good or ill. Get the incentives wrong and you buy the appearance of quality; get them right and you buy the thing itself.

Core concepts

Quality and its dimensions. Health care quality is multidimensional: care that is safe, effective, patient-centred, timely, efficient, and equitable is the widely used shorthand. No single number captures it, which is why measurement is contested and why a payer must be explicit about which dimension a given indicator is standing in for.

The Donabedian triad. The foundational framework for measuring quality is the Donabedian model, which distinguishes structure (the physical and organizational resources — staffing, equipment, accreditation), process (what is actually done to patients — whether antibiotics were given on time), and outcome (what happens to health — mortality, infection, function). Structure is easy to measure but weakly linked to results; outcomes are what we care about but are slow, noisy, and confounded by case mix; process measures sit between, being actionable and quick but only as good as the evidence linking them to outcomes. Most measurement schemes mix all three.

Harm, error, and adverse events. An adverse event is an injury caused by medical management rather than the underlying disease; when it results from a mistake it is a medical error, and harm caused by the care itself is iatrogenesis. Not all adverse events are preventable and not all errors reach the patient, but the preventable, harmful subset is where the economics concentrates. A never event is a serious, largely preventable safety incident — wrong-site surgery, a retained instrument — that should never occur if defined safeguards are followed, and which many payers refuse to pay for.

Healthcare-associated infection. A hospital-acquired infection — one acquired during care rather than present on admission — is the archetypal measurable, costly, and partly preventable harm. Central-line-associated bloodstream infections, catheter-associated urinary infections, and surgical-site infections each carry a well-studied excess cost in bed-days and treatment, which makes them the natural proving ground for the business case for safety.

The cost of poor quality and the business case for safety. The business case for safety asks whether an investment in prevention pays for itself in avoided failure costs. It is genuinely uncomfortable, because the party that funds prevention is not always the party that reaps the saving, and because a prevented infection is an invisible non-event that no budget line celebrates. Establishing the case requires knowing the attributable cost of the harm, the effectiveness of the intervention, and — crucially — whose budget each falls on.

Paying for quality. Pay-for-performance (P4P) ties a portion of provider payment to measured quality or safety. Value-based purchasing is the broader move to buy outcomes and quality rather than mere activity. Both sit on top of the underlying payment mechanism (capitation, fee-for-service, diagnosis-related groups), which belongs to Chapter 3.1 — Health Systems; here the concern is the quality overlay and its behavioural effects.

Goodhart's law and gaming. The central hazard of paying for measured quality is Goodhart's law: when a measure becomes a target, it ceases to be a good measure. Providers optimize the indicator rather than the underlying goal — teaching to the measure, up-coding, or avoiding the sickest patients — so a scheme can raise the score while leaving real quality untouched or worse.

Root cause analysis. After a serious incident, root cause analysis is the structured investigation that looks past the individual to the system failures — the missing checklist, the understaffed shift, the confusing label. Its economic value is that fixing a systemic cause prevents a class of future harms, not just a repeat of one.

Best practices

  1. Cost the harm before you argue about the fix. You cannot make a business case for safety without an attributable cost — the excess bed-days, treatment, readmission, and litigation caused by a specific harm, over and above the cost of the same patient without it. Use local data where you can and published attributable-cost studies where you cannot, and be explicit about which costs are cash-releasing and which merely free a bed for another patient. A prevented infection that only frees capacity, without cutting spend, is still valuable — but say so honestly rather than claiming savings that never appear in the ledger.

  2. Name whose budget the harm and the fix fall on. The business case collapses when the ward that must fund the prevention is not the budget that carries the cost of the harm. A safety investment that pays back at the level of the whole system may look like pure cost to a departmental manager whose savings leak to the insurer or the next provider in the pathway. Identify the split early, and design the funding flow — a pooled safety fund, a shared-savings arrangement — so that the investor is not left worse off for doing the right thing.

  3. Anchor measurement in the Donabedian triad and mix the three deliberately. Do not rely on structure measures because they are convenient, nor demand outcomes alone because they are slow and confounded. Pair a small number of high-value outcome measures with the process measures that are known to drive them, and use structure measures only as enabling conditions. Every process measure you reward should have credible evidence linking it to an outcome patients care about — otherwise you are paying for compliance, not health.

  4. Prefer preventing harm to detecting and paying for it later. The cheapest adverse event is the one that never happens, and the highest-return safety investments are usually cheap, standardized practices — hand hygiene, insertion checklists, surgical time-outs, safe-staffing thresholds — rather than expensive technology. Fund the reliable delivery of known safe practices before you buy novel monitoring. Reliability, not novelty, is what converts an evidence-based practice into an avoided harm.

  5. Use penalties for never events sparingly and precisely. Refusing to pay for a genuinely never-should-happen event — a retained instrument, wrong-site surgery — sends a clear signal and aligns with public expectations. But penalties work only when the event is truly preventable, unambiguously defined, and reliably detected; applied to harms that are only partly avoidable, they punish case mix and bad luck, drive under-reporting, and corrode the honest incident-reporting culture that safety depends on. Reserve hard penalties for the small, defensible list.

  6. Design pay-for-performance to resist gaming from the start. Assume providers will optimize whatever you measure, and stress-test each indicator against Goodhart's law before you launch. Watch for the three classic distortions: teaching to the measure (effort shifts from unmeasured to measured care), up-coding (recording sicker patients or ticking boxes without doing the work), and cream-skimming or patient selection (avoiding the complex, high-risk patients who threaten the score). Use basket measures rather than single indicators, audit a sample of the underlying records, and rotate or retire measures that have saturated.

  7. Adjust for case mix, or you will punish those who treat the sickest. Raw outcome comparisons reward providers who select easy patients and penalize those who take the hardest. Risk-adjust outcome-based rewards and penalties so that a hospital serving deprived, multimorbid populations is not marked down for its case mix — otherwise the scheme creates an incentive to avoid exactly those patients, worsening equity (see Chapter 3.4 — Equity). Publish the adjustment method; an opaque model destroys the trust the scheme needs.

  8. Protect reporting culture as an economic asset. Learning from harm requires that staff report near-misses and errors without fear, because the incidents you never hear about are the ones you cannot prevent. Incentive schemes that penalize reported harm can suppress the very data that drives improvement, so separate the learning system from the payment system, and treat a rising reported-incident count during a safety drive as a sign of trust, not decline. A blame-free root cause analysis that fixes the system prevents a whole class of future costs.

  9. Size the incentive to change behaviour without distorting it. Pay-for-performance payments that are too small to notice change nothing; payments too large tempt gaming and destabilize provider finances when a single metric swings. Evidence on P4P is mixed — effects are often modest, sometimes short-lived, and occasionally perverse — so treat any scheme as a hypothesis to be evaluated, not a settled instrument. Start modest, evaluate rigorously (see Chapter 2.3 — Health Econometrics), and be willing to stop.

  10. Evaluate the quality scheme as you would any intervention. A pay-for-performance or value-based-purchasing programme is itself a health investment with costs — measurement, reporting, administration, and the clinical time diverted to it — and it should clear the same bar as the care it governs. Build in a counterfactual, watch for unintended effects on unmeasured care, and count the administrative burden as a real cost. If the scheme cannot demonstrate that it buys more health than the same money spent directly on care, retire it.

Questions to discuss with your team

  1. Which harms cost us the most, and do we actually know their attributable cost or are we guessing? Most organizations can name their high-profile incidents but cannot say what an avoidable harm truly costs them, because the excess is buried across bed-days, readmissions, and litigation held on different budgets. The honest starting point is to pick two or three high-volume harms — a healthcare-associated infection, a fall, a pressure injury — and assemble the real local cost, distinguishing cash you would release from capacity you would free. Expect the exercise to be harder than it sounds and the numbers to be contested; that difficulty is itself the finding. A good answer names the data you do not yet have and commits to closing the gap, rather than reaching for a figure from another country's study and treating it as your own.

  2. If we pay for a quality measure, how will people game it — and can we live with that? Every measure you attach money to will be optimized, and the question is not whether gaming will happen but which form you can tolerate. Walk through each candidate indicator and ask who is disadvantaged if a provider chases the score: will the sickest patients be avoided, will unmeasured care be neglected, will records be up-coded, will honest reporting be suppressed? An honest discussion admits that some distortion is the price of any incentive and decides deliberately where to accept it, where to guard against it with audit and risk adjustment, and where the gaming risk is so severe that the measure should not carry money at all. The failure mode is a team that believes its own indicator is gaming-proof.

  3. Is our safety investment justified by the value case, or only by the value case at system level that no local budget will fund? Safety investments frequently pay back for the whole system while leaving the department that must fund them out of pocket, because the savings land in another budget or another organization. The team needs to be candid about whether a proposed investment survives at the level of the budget holder who has to say yes, and if it does not, what mechanism — a pooled fund, shared savings, a mandate from the payer — could make the case whole. This is where safety economics meets organizational reality: a genuinely cost-saving intervention can still fail to happen because no one who benefits is the one who pays. A serious answer traces the money, not just the health.

  4. For the outcomes we care about, are we measuring the processes that actually drive them, or just the numbers that are easy to collect? The Donabedian triad tempts every organization towards its convenient corner: structure measures because accreditation already counts them, or outcome measures because they sound like the thing that matters, when the actionable middle — the process measures known to move outcomes — is where improvement lives. The discussion should take two or three outcomes the organization is judged on and ask, for each, whether the processes being tracked have credible evidence linking them to that outcome, or whether they are proxies of convenience. It should also confront the opposite error: rewarding rare outcomes on single wards, where the numbers are too noisy to steer by and a good or bad quarter is mostly chance. An honest answer distinguishes the handful of measures worth attaching consequences to from the larger set worth watching but not rewarding, and admits where the evidence link is assumed rather than known.

  5. Which of our harms are truly "never events", and where would a hard penalty do more harm than good? The appeal of non-payment for never events is moral clarity, but the instrument only works where the event is genuinely preventable, unambiguously defined, and reliably detected — and most harms fail at least one of those tests. The team should sort its serious incidents into the small set that meets all three conditions and the larger set that is only partly avoidable, where a penalty would punish case mix and chance, drive attribution games, and chill the reporting the organization depends on. This is uncomfortable because it means admitting that some harms cannot be penalized into nonexistence without collateral damage. A good answer names the defensible penalty list explicitly and chooses a gentler instrument — a risk-adjusted basket, a shared-savings incentive — for everything outside it.

  6. How would we know if our quality scheme is buying health rather than just buying scores? Every pay-for-performance or value-based-purchasing programme is itself an intervention with real costs — measurement, reporting, administration, and the clinical time diverted into it — yet schemes are routinely launched without a counterfactual or a stop rule, and then run indefinitely on faith. The team should ask what evidence would persuade it that the scheme is working, what would persuade it the scheme is failing or distorting unmeasured care, and whether either kind of evidence is actually being collected. It should be candid that the honest answer might be "we cannot currently tell", which is itself a reason to build evaluation in before expanding. A serious response commits to a pre-specified comparison, watches for effects on the care nobody is measuring, and is willing to retire a scheme that cannot show it earns its keep.

In practice: a health economics example

The Larkmoor Hospitals Group is a fictional not-for-profit network of four acute hospitals in the high-income, social-insurance funded Republic of Ostheim, paid largely through diagnosis-related groups by competing sickness funds. Its new director of quality is asked to justify a proposed group-wide programme to cut central-line-associated bloodstream infections (CLABSIs) — a serious hospital-acquired infection — after a cluster of cases in intensive care. The proposed programme is unglamorous: a standardized insertion checklist, a stocked insertion cart, staff training, daily review of line necessity, and an infection-prevention nurse to sustain it. Its annual cost is modest and mostly staff time.

The director begins by costing the harm. Rather than importing a headline figure, the team assembles Ostheim data: each CLABSI adds a well-documented run of intensive-care and ward bed-days, extra antibiotics and diagnostics, and a raised mortality risk. Under the diagnosis-related group system, the sickness fund pays a fixed amount for the admission regardless of the infection, so the excess bed-days are largely a cost to the hospital, not the payer — which, unusually, means the business case is strong precisely because the provider bears the harm. The team separates cash-releasing savings (agency staff, drugs) from freed capacity (ICU beds that can now take other funded admissions), and is careful not to double-count.

The economics are favourable: the programme's cost is a fraction of the avoided failure cost even under conservative assumptions about how many infections it prevents, and a one-way sensitivity analysis shows the case holds even if effectiveness is half the published estimate. But the director resists three temptations. She does not promise the finance director a cash saving where the benefit is really freed ICU capacity, because a saving that never appears in the ledger would discredit the next safety case. She structures the Donabedian model measurement to track process (checklist compliance, line-days) alongside the outcome (infection rate per 1,000 line-days), knowing the outcome is too rare on a single ward to move quickly. And she keeps incident reporting separate from any performance conversation, so that a rise in reported line concerns during the drive is read as vigilance, not failure.

One board member proposes going further: docking a ward's budget for every CLABSI, treating it as a never event. The director advises against it. CLABSIs are largely preventable but not wholly so; a hard penalty would fall partly on case mix and chance, would tempt staff to attribute infections to admission rather than the line, and would poison the reporting culture the programme depends on. Ostheim's payer, she notes, is separately piloting a modest value-based purchasing adjustment that rewards low risk-adjusted infection rates across a basket of measures — a gentler, harder-to-game instrument than a single-event penalty. The board funds the programme as a system investment, with a commitment to evaluate it against the pre-programme trend after eighteen months rather than declaring victory on the first good quarter.

Four sector lenses

Startup

A digital health start-up selling infection-surveillance or safety-monitoring software lives or dies on the strength of its business case for safety, because its buyer is a hospital asking "what harm will this prevent, and is the avoided cost bigger than your licence fee?" The honest start-up quantifies avoided harm in the buyer's own cost terms and distinguishes cash from capacity, rather than citing a foreign attributable-cost study as if it transferred. Its danger is selling detection when the buyer's real gap is reliable action on what is already known — a dashboard that surfaces more alerts without changing behaviour adds cost, not safety. Early evidence is thin, so credible pilots with a counterfactual beat impressive-sounding claims.

Small business

A small but established provider — a general-practice partnership, a single-site clinic, a care home, or a specialist device supplier — carries the harm and funds the prevention within the same modest budget, which sharpens the business case but leaves little slack to absorb a bad quarter. Unlike a start-up, it is not chasing growth on thin evidence; it runs a steady service and mostly needs to embed a handful of known safe practices — hand hygiene, medication reconciliation, pressure-injury checks — reliably rather than to invent anything new. Its constraint is capacity and scale: it cannot risk-adjust its own outcomes across sites, and a single serious incident can swamp its small numbers, so raw comparison against larger peers is unfair and internally misleading. It usually enters pay-for-performance and never-event schemes as a price-taker, so it should focus on the few processes it can deliver consistently and lean on external benchmarks and shared learning rather than building measurement machinery it cannot afford.

Enterprise

A large hospital group or insurer has the scale to build the attributable-cost data, run risk-adjusted comparisons across sites, and operate a serious pay-for-performance or value-based-purchasing scheme — and the scale to do real damage if it gets the incentives wrong. Its central challenge is the split-budget problem within its own walls: the ward that funds prevention, the finance function that counts savings, and the pathway that inherits the freed capacity are different accountabilities, so it must engineer internal funding flows that reward the investor. At enterprise scale, Goodhart's law bites hardest, because a group-wide metric tied to money will be optimized across dozens of sites; basket measures, audit, and case-mix adjustment are non-negotiable.

Government

A ministry or national payer sets the rules of the game for everyone — the never-event list, the value-based-purchasing formula, the penalties for healthcare-associated infection — and its choices propagate system-wide. The United States Medicare programme, for example, penalizes hospitals in the worst-performing quartile for hospital-acquired conditions and adjusts payment through value-based purchasing; England has used Commissioning for Quality and Innovation (CQUIN) schemes and a national never-events list with non-payment. Government's unique risk is that a national incentive, however well-meant, is a blunt instrument applied to heterogeneous providers, so it must risk-adjust to protect those serving the sickest and deprived populations (see Chapter 3.4 — Equity), evaluate rigorously, and be willing to retire schemes that buy scores rather than health.

Common failure modes

  • Treating safety as a pure cost. Framing prevention as spending that competes with productivity, ignoring that harm consumes far more resource than prevention. Fix: build and publish the attributable-cost case so safety is seen as released capacity.

  • Claiming cash savings that never materialize. Promising the finance director money when the real benefit is freed capacity, then losing credibility when the ledger does not change. Fix: separate cash-releasing from capacity-freeing benefits explicitly.

  • Paying for a single metric. Attaching money to one indicator, which is then optimized at the expense of everything unmeasured. Fix: use baskets of measures, audit the underlying records, and retire saturated metrics.

  • Ignoring case mix. Rewarding or penalizing raw outcomes, which punishes providers who treat the sickest and rewards patient selection. Fix: risk-adjust transparently before money changes hands.

  • Penalizing honest reporting. Docking budgets for reported harm, which suppresses the incident data that improvement depends on. Fix: firewall the learning system from the payment system.

  • Confusing process compliance with quality. Rewarding a process measure with no proven link to outcomes, buying box-ticking rather than health. Fix: require evidence linking each rewarded process to an outcome patients value.

  • Launching a scheme you never evaluate. Running pay-for-performance indefinitely without a counterfactual, unable to say whether it buys health. Fix: treat every scheme as a hypothesis with a built-in evaluation and a stop rule.

Maturity model

Dimension Initiate Develop Standardize Manage Orchestrate
Costing harm Harm cost unknown; safety argued on principle alone Headline figures borrowed from external studies Local attributable costs assembled for key harms, cash vs capacity split Costs maintained, updated, and used to prioritize investment Attributable costs shared across the pathway and used to negotiate funding flows between organizations
Measurement Ad hoc structure measures; outcomes untracked Some process and outcome measures collected but unlinked Donabedian mix with evidence-linked process measures and risk adjustment Measures rotated as they saturate; case mix and equity routinely adjusted Measurement aligned across sites and partners; benchmarks and definitions shared system-wide
Paying for quality Payment tied only to activity Single quality metric attached to money Basket measures, audited, with gaming risks assessed before launch Incentives sized by evidence and evaluated against a counterfactual Incentives coordinated across payers and providers, retired when they stop adding value
Reporting culture Blame culture; harm under-reported Reporting exists but feared to affect budgets Learning system firewalled from payment; root cause analysis routine Rising reports read as trust; systemic fixes tracked to closure Learning pooled across organizations so a fix in one prevents a class of harm in all
Business case ownership Split budgets block investment Split acknowledged but unaddressed Funding flows engineered so the investor is rewarded Investment prioritized on maintained cost and effectiveness data System-level value realized through pooled funds and shared-savings arrangements across partners

Checklist

  • Attributable local cost estimated for your two or three highest-volume avoidable harms, with cash-releasing and capacity-freeing benefits distinguished.
  • The budget that funds each prevention and the budget that carries each harm identified, and any split addressed by a funding mechanism.
  • Every rewarded process measure backed by evidence linking it to an outcome patients care about.
  • Each money-carrying indicator stress-tested against gaming (teaching to the measure, up-coding, patient selection) before launch.
  • Outcome-based rewards and penalties risk-adjusted for case mix, with the method published.
  • Never-event penalties limited to truly preventable, unambiguously defined, reliably detected events.
  • Incident-reporting and learning systems firewalled from the payment system.
  • Each quality or pay-for-performance scheme launched with a counterfactual, an evaluation plan, and a stop rule.
  • Administrative and clinical-time costs of the measurement scheme counted as real costs.

Key sources

  • Donabedian, A. — "Evaluating the Quality of Medical Care" — the structure–process–outcome framework underpinning quality measurement.
  • World Health Organization — patient safety and Global Patient Safety Action Plan — the international framing of avoidable harm as a major, largely preventable burden.
  • Institute of Medicine — To Err Is Human and Crossing the Quality Chasm — the reports that established the systems view of error and the six dimensions of quality.
  • US Centers for Medicare & Medicaid Services — Hospital Value-Based Purchasing and Hospital-Acquired Condition Reduction Program — national exemplars of paying and penalising for quality and safety.
  • NHS England — Never Events policy and framework, and Commissioning for Quality and Innovation (CQUIN) — English exemplars of non-payment for never events and quality-linked payment.
  • OECD — Health at a Glance and work on the economics of patient safety — international evidence on the cost of poor quality and avoidable harm.

References

  1. Patient safety — Wikipedia — https://en.wikipedia.org/wiki/Patient_safety
  2. Health care quality — Wikipedia — https://en.wikipedia.org/wiki/Health_care_quality
  3. Donabedian model — Wikipedia — https://en.wikipedia.org/wiki/Donabedian_model
  4. Adverse event — Wikipedia — https://en.wikipedia.org/wiki/Adverse_event
  5. Medical error — Wikipedia — https://en.wikipedia.org/wiki/Medical_error
  6. Iatrogenesis — Wikipedia — https://en.wikipedia.org/wiki/Iatrogenesis
  7. Never event — Wikipedia — https://en.wikipedia.org/wiki/Never_event
  8. Hospital-acquired infection — Wikipedia — https://en.wikipedia.org/wiki/Hospital-acquired_infection
  9. Pay for performance (healthcare) — Wikipedia — https://en.wikipedia.org/wiki/Pay_for_performance_(healthcare)
  10. Value-based purchasing — Wikipedia — https://en.wikipedia.org/wiki/Value-based_purchasing
  11. Goodhart's law — Wikipedia — https://en.wikipedia.org/wiki/Goodhart%27s_law
  12. Root cause analysis — Wikipedia — https://en.wikipedia.org/wiki/Root_cause_analysis