Why this matters in health economics

Every encounter a health system records — every prescription, image, laboratory result, and admission — creates data twice over: once as the working record of an individual's care, and again as a potential input to something larger. Combined and analysed, the same records can answer questions no trial will ever fund: which treatments work in real populations, where safety signals hide, which services fail which communities, how a pandemic moves. Health systems sit on decades of this asset, and most of them realize a small fraction of its value — not because the science is missing, but because the data is fragmented across incompatible systems, of uneven quality, locked behind access processes measured in years, or guarded by institutions with no incentive to share.

The waste is quantifiable in kind if not always in number: research not done, duplicated data collection, decisions made on intuition where evidence existed but could not be assembled, and analysts — among the scarcest labour in the system — spending most of their time finding, cleaning, and reconciling data rather than analysing it. Meanwhile the asset decays: data uncurated is data depreciating, as meanings drift, documentation is lost, and the staff who understood a dataset move on.

Against the value stands a real liability. Health data is among the most sensitive information that exists about a person, and its misuse — a breach, a sale that surprises the public, a use that feels like betrayal — carries costs far beyond any fine: it spends the public trust that every future use depends on. England's care.data programme (abandoned in 2016) remains the canonical lesson: a technically reasonable plan to pool general-practice records for research and planning collapsed against public and professional concern about consent and commercial access, and its failure set back legitimate data use for years. For a director, this is the defining trade-off of the chapter: the value of health data is unlocked by access, and the licence for access is trust — so the economics of data are inseparable from the governance of data.

Core concepts

Health data is economically peculiar. It is non-rival — one team's use does not consume it, so the same dataset can serve unlimited users at near-zero marginal cost — which makes it behave like a public good within its access boundary and means under-use is pure waste. But its fixed costs are heavy and continuous: collection, curation, standardization, documentation, and stewardship must be paid whether one study uses the data or a thousand do. And its value is combinatorial — records gain value when linked across settings and time, which is precisely what fragmentation across institutions and vendors prevents. Secondary use — analysing data collected for care to serve research, planning, and improvement — is where most of the unrealized value lies.

The architecture that holds the data embodies an economic strategy. The data warehouse is the classical model: data from source systems is cleaned, standardized, and loaded into a central store built for analysis — high curation cost up front, high usability after. The data lake inverts the bargain: raw data of any shape is stored cheaply now and interpreted later — which defers the curation cost but does not remove it, and a lake without stewardship becomes the proverbial swamp, where data is technically present and practically unusable. The data mesh is the newest strategy, and it is organizational rather than technological: instead of one central team owning all data, each domain — the laboratory, the emergency department, the mental-health service — owns and publishes its data as a product, to common standards, with the platform team providing shared infrastructure and the governance federated. Its economic logic is that the people who generate data hold the context needed to make it meaningful, and a central team is a bottleneck that scales badly; its economic risk is that it requires every domain to fund and staff data stewardship — a real, distributed cost that a central model at least makes visible in one place.

No architecture beats bad inputs. Data quality — completeness, accuracy, timeliness, consistency — is an economic quantity: poor-quality data taxes every downstream use, and the tax compounds, since analysts re-clean the same defects study after study. Quality is made at the point of care, where the clinician coding a diagnosis has weak incentives to serve a future analyst, so improving it is an incentive-design problem as much as a technical one (the general measurement machinery is Chapter 2.3 — Health Econometrics, which owns what bad data does to causal claims).

Access is where value is realized or forfeited, and access has its own cost structure. Because the marginal cost of use is near zero, the efficient policy is wide access at low friction — but each access carries privacy risk, so systems impose governance: approval processes, de-identification (removing or masking identifiers, always a spectrum of residual re-identification risk rather than a binary), and increasingly the trusted research environment model, in which analysts come to the data inside a secure enclave and only vetted results leave — replacing the riskier habit of shipping copies of datasets outward. Techniques such as differential privacy (adding calibrated statistical noise so no individual's presence can be inferred) and federated learning (training models across distributed datasets without pooling the records) shift the privacy–utility frontier outward. The economic point is that access friction is a real price: a two-year approval process rations data not by value of use but by applicants' stamina, and the studies it deters are an invisible cost no ledger records.

The rules around all this are data governance: who may use what data for which purposes, decided how, and accountable to whom. Regulation sets the floor — the General Data Protection Regulation (GDPR) in Europe, sector rules such as the United States' HIPAA elsewhere — and new structures are being built above it: the European Health Data Space (regulation adopted 2025) creates a legal framework for both primary and secondary use of health data across the European Union, and national access bodies such as Finland's Findata show what a single front door for permits and secure access looks like in practice. The FAIR data principles — findable, accessible, interoperable, reusable — summarize the stewardship standard that makes any of it work. The plumbing that moves the data — interoperability standards such as HL7 FHIR — belongs to Chapter 5.4 — Software Engineering Health Economics; the models trained on the data to Chapter 5.3 — AI Health Economics; and the public-trust dynamics to the communication economics of Chapter 4.4 — Social Media and Health Communication Economics.

Best practices

  1. Treat data as an asset with an owner, a budget, and a depreciation schedule. Inventory the significant data assets you hold, name a steward for each, and fund curation and documentation as standing costs — not as residue of projects. An unstewarded dataset is a depreciating one, and the cheapest time to capture context is while the people who created the data still remember it.

  2. Choose architecture by organizational shape, not fashion. Centralize (warehouse or governed lake) when analytical needs are shared, domains are weak in data capability, and one team can hold the context; federate toward a mesh when domains are strong, needs are diverse, and the central team has become the queue everyone waits in. Either way, the curation cost is conserved — the choice is who pays it and where the bottleneck sits. Migrating architecture is expensive; do it for a named economic failure (a bottleneck, a swamp), not a conference talk.

  3. Never build a lake without a stewardship budget. Cheap storage makes it easy to defer curation forever; deferred curation is how lakes become swamps in which every use starts from raw archaeology. Fund cataloguing, metadata, and quality monitoring as part of the platform, and measure the lake by data actually used, not terabytes held.

  4. If you adopt mesh principles, fund the domains to be publishers. Data-as-a-product means each domain must staff and pay for the stewardship of what it publishes — a real cost that central models concentrate and mesh models distribute. Set common standards centrally (identifiers, metadata, quality thresholds, access rules), fund domain data roles explicitly, and accept that a mesh without funded domain ownership is just fragmentation with better branding.

  5. Fix data quality at the source, with incentives, not exhortation. Downstream cleaning re-pays the same tax forever; upstream fixes pay once. Make quality visible where it is created (feedback to the coding clinician or clerk), make it matter (quality-sensitive payment and audit where appropriate), and design forms and systems so the easy path is the accurate one — this is choice architecture applied to documentation (see Chapter 4.1 — Behavioural Economics).

  6. Price access friction as a cost, and drive it down deliberately. Measure the full journey a legitimate analyst faces — discovery, permit, extraction or enclave entry — in months and staff-hours, and treat reductions as value creation. A single front door (one application, one decision-maker, published timelines, standard contracts) is the reform with the highest yield; Finland's Findata model is the reference pattern.

  7. Default to analysts visiting data, not data visiting analysts. Trusted research environments — secure enclaves with vetted egress — dominate dataset shipping on both risk and auditability, and they make wide access compatible with strong control. Reserve extracts for the cases enclaves genuinely cannot serve, and treat every copied dataset in the wild as standing liability.

  8. Match the privacy technique to the use, and be honest about residual risk. De-identification is a spectrum, not a switch; differential privacy buys mathematical guarantees at a utility price; federated learning avoids pooling but not all leakage. Choose per use-case, document the residual risk honestly, and never market "anonymized" as "risk-free" — overclaiming is how trust is lost when the inevitable exception surfaces.

  9. Earn the social licence before you need it. The public's tolerance for data use is specific: strong for care and public-benefit research, fragile for commercial access and anything that surprises. Be transparent about who uses data for what (publish the register), involve patients and the public in access policy, offer genuine choice where the law and the use-case allow, and remember care.data: consent processes are cheaper than collapses.

  10. Set the terms of commercial access in public, in advance. Industry use of health data can be legitimate and valuable — trials, safety surveillance, product development — and it is also where trust is most flammable. Decide openly what commercial uses are permitted, what the public system charges or receives (money, capability, guaranteed access to resulting products), and what is off the table, before the deal, not after the headline.

  11. Charge for access on marginal cost plus stewardship, not monopoly rent. A public data holder that prices access to maximize revenue rations a near-zero-marginal-cost asset and forfeits public value; one that charges nothing starves stewardship. Cost-recovery pricing with waivers for public-interest research funds the asset without strangling its use.

  12. Count the value you create, because no one else will. Data programmes are perpetually vulnerable at budget time for the same reason communication is: their product is other people's results. Track and publish what access enabled — studies, service changes, safety findings, avoided duplicate collection — and put the portfolio through the same value scrutiny as any investment (Chapter 2.1 — Economic Evaluation), so the asset's case is evidence rather than assertion.

Questions to discuss with your team

  1. What data do we actually hold, who stewards each asset, and what is it costing us to keep it usable — or to let it rot? Most organizations cannot produce the inventory: datasets accumulate in systems, projects, and shared drives, with stewardship assigned to nobody and documentation living in departed employees' memories. Build the register — the significant assets, their owners, their quality, their documentation, their use — and attach the standing cost of proper curation to each, alongside the depreciation now under way where there is none. The tension is that stewardship is a permanent unglamorous cost with diffuse benefits, exactly the line every budget round wants to cut, while the losses from unstewarded data (re-collection, unusable history, analysts as archaeologists) never appear on any ledger. An honest answer produces the inventory with named stewards, funds curation as infrastructure rather than project residue, and admits which assets have decayed past saving.

  2. Where does an analyst's time actually go — and is our architecture the reason? The scarcest resource in the data system is skilled analytical labour, and in most health organizations the majority of it is spent finding, accessing, cleaning, and reconciling data rather than analysing it. Measure the split honestly for a few recent projects, then trace the causes: fragmented sources, absent documentation, quality defects re-cleaned for the hundredth time, a central data team that has become a queue. This is the diagnostic that should drive architecture — warehouse, lake, or mesh — because each is a different answer to where the bottleneck sits and who holds context. The tension is that architecture debates are conducted in vendor vocabulary and settled by fashion, while the actual question is organizational: who owns which data, who pays for its curation, and where the queue forms. An honest answer states the current analyst-time split, names the binding bottleneck, and chooses architecture as a remedy for that bottleneck — with the migration's cost justified by the labour it frees.

  3. How long does it take a legitimate user to get from question to data, and what is that delay costing? Walk the full journey concretely — a researcher, a service planner, a quality team — through discovery, application, approval, contracting, and access, and count the months and the staff-hours on both sides. Then ask what never happens because of it: the studies not attempted, the decisions taken unevidenced, the analysts who route around the system or leave. The tension is real on both sides — every relaxation of friction is a governance decision with risk attached, and every layer of process was added after some incident — but friction rations by stamina, not by value of use, and its costs are invisible while its protections are auditable. An honest answer measures the end-to-end time, separates the delay that buys genuine protection from the delay that is queueing and duplication, sets a published target, and builds the single front door that makes the target achievable.

  4. What is our data quality where it is made, and what incentive does anyone have to improve it? Quality problems discovered in analysis were created months earlier at a keyboard in a clinic, by someone whose job that day was care, not coding. Examine your key datasets for completeness and accuracy, then follow the defects upstream: does the clinician or clerk who records the data ever see its downstream use, get feedback on its quality, or gain anything from improving it? Does the system make accurate recording the easy path or the heroic one? The tension is that quality-improvement pressure lands on the busiest people in the system and can backfire into gaming if tied crudely to payment (the measurement pathologies of Chapter 3.11 — Quality and Safety Economics apply squarely). An honest answer measures quality at source, closes the feedback loop so data creators see what their data becomes, redesigns the capture path before blaming the people on it, and reserves payment incentives for what audit can actually verify.

  5. Whom would our data uses surprise, and what is our social licence actually good for? The care.data collapse was not caused by unlawful processing; it was caused by uses the public had not been told about in terms they would have accepted, discovered in the worst possible way. Inventory your current and planned data uses — internal analytics, research access, any commercial arrangements — and test each against the surprise standard: if this appeared on a front page tomorrow, would patients feel informed or betrayed? Probe the specifics that history says are flammable: commercial access terms, data leaving your custody, uses that look like surveillance or profiling of the very communities whose trust is thinnest. The tension is that full transparency feels risky to programmes that grew up quietly, and public involvement takes time that delivery timelines resent — but the alternative is borrowing against a licence you may not hold, at care.data interest rates. An honest answer publishes the use register, involves patients and the public in access policy before the controversial case arrives, sets commercial terms in public, and treats the licence as a measured asset with a trajectory, not an assumption.

  6. Are we set up to prove this programme's value, or are we asking the budget to take it on faith? Data platforms, stewardship teams, and access services produce their value in other people's results, which is why they are perennially first against the wall in a savings round — the costs are concentrated and visible, the benefits diffuse and credited elsewhere. Ask what you currently record about what access enabled: studies completed, service changes made, safety signals found, collections not duplicated, analyst-hours saved by curation done once instead of many times. Then ask the harder portfolio question: which of your data investments are earning, which are speculative, and which are monuments. The tension is that attribution is genuinely hard — data is one input among many to any outcome — and overclaiming corrodes credibility as surely as silence starves it. An honest answer builds outcome-tracking into the access process itself, publishes an annual account of realized value with the attribution caveats stated plainly, and is willing to retire the platform component or dataset whose use never materialized.

In practice: a health economics example

A fictional Nordic country's national health-data authority is created to fix a paradox: the country has decades of rich, linkable health records — a national identifier, universal coverage, digitized registries — and yet researchers routinely wait two years for access, each of twenty regions answers permit requests differently, and two regions have started building their own competing "data lakes" with no stewardship plan. A pharmaceutical company recently abandoned a post-market safety study because the permits could not be assembled; a university team took a grant to a neighbouring country instead. The value is provably there; the system is forfeiting it.

The authority's economists begin by measuring the friction as a cost. They trace twelve recent access attempts end to end — discovery, application, twenty separate regional decisions, contracting, extraction — and cost the journey in months (median twenty-six) and staff-hours on all sides, then list what was abandoned along the way. Presented this way, access reform stops being an administrative nicety and becomes an investment case: the single front door — one application, one legal decision-maker, published timelines — is projected to release more research value per krona than any new dataset the country could buy. The design borrows deliberately from Finland's Findata: a national permit authority plus a trusted research environment, so analysts come to the data and vetted results leave, replacing the shipped extracts that currently sit unaudited on university servers.

The architecture decision is treated as an organizational economics question, not a technology procurement. A fully central national warehouse is rejected: the regions hold the clinical context, and a single national team would become the new bottleneck. Pure regional autonomy is also rejected: twenty lakes with twenty metadata standards is the swamp, twenty times. The authority lands on mesh principles under federated governance — each region and registry publishes its data as a product, to national standards for identifiers, metadata, quality reporting, and access — with the national platform providing the shared enclave, catalogue, and permit service. The honest cost of this choice is made explicit: every region must fund data-steward roles, and the authority's budget includes transitional funding to make domain ownership real rather than nominal, because a mesh no one staffs is fragmentation with better branding.

Trust is engineered as deliberately as the platform. The authority studies the care.data failure explicitly and does the opposite: the full register of data uses is public from day one; a citizens' panel sits inside the access-policy process with teeth, not tokens; commercial access is permitted under terms set in public in advance — cost-recovery charging, enclave-only access, published summaries of every approved use, and a share of value returning to the system in capability and guaranteed research access. When a tabloid runs a "your records for sale" story in year two, the programme survives it — not because the story is kind, but because every fact it alleges is already on the authority's website, with the terms attached; the pre-built transparency does the rebuttal (the communication economics are Chapter 4.4 — Social Media and Health Communication Economics).

Three years in, the authority publishes its accounting: median time from application to enclave access down from twenty-six months to four; studies enabled, service analyses delivered, and one genuine safety signal found and acted on; regional data-quality dashboards moving because clinicians now see what their coding becomes. The evaluation is candid about attribution limits and about one failure — a region whose "data products" remain undocumented and unused, where domain funding was cut — which the authority uses to argue, successfully, that stewardship is the asset's maintenance bill, not an optional extra.

Four sector lenses

Startup

A health-data start-up — an analytics product, a curation tool, a federated-learning platform — lives on data it does not own, so its existential questions are access and trust: what legitimate basis it has for the data that feeds it, how it gets through data holders' governance at a survivable pace, and whether one privacy incident ends the company. The temptation is to treat governance as friction to be hacked; the durable play is the opposite — build for enclaves, auditability, and privacy-preserving techniques from the start, because that is what the buyers' governance will demand, and being easy to approve is a moat. Its business-model risk is the data-broker shadow: models that resell or monetize health data in ways that would surprise patients are one headline from taking the whole sector's licence down with them.

Small business

A small provider — a clinic group, a pharmacy, a care home operator — is a data creator more than a data user: its records feed national registries, payer datasets, and research it will never see. Its economics are about burden and return: data quality is made or broken at its keyboards, yet it is rarely paid, thanked, or fed back for the quality it creates, and every new reporting requirement is unfunded labour. The sensible posture is to standardize ruthlessly (use the standard codes and systems, resist bespoke extracts), to demand feedback loops and burden-reduction in return for reporting — a fair exchange the wider system should be embarrassed to refuse — and to treat its own modest data as an operational asset for scheduling, stock, and quality rather than waiting for a platform to arrive.

Enterprise

A large provider group or insurer holds data at the scale where the asset economics become real: longitudinal records across millions of lives, enough volume for credible analytics, and enough incidents to know the liability side too. The mature enterprise runs data as a governed portfolio — an inventory with stewards, a funded platform with a published access process for its own teams, quality managed at source, and architecture chosen for its organizational shape rather than its conference calendar. Its distinctive temptations are hoarding (guarding data against the system it serves, forfeiting combinatorial value) and quiet monetization (commercial deals that patients discover from journalists). The distinctive discipline is to treat internal access friction as seriously as external: an insurer whose own actuaries wait months for data has built a bottleneck, not a governance function.

Government

A ministry or national agency owns the layer no one else can: the identifiers and standards that make linkage possible, the legal framework for secondary use, the national access institutions, and the terms on which the country's data asset serves research and industry. The reference patterns now exist — Finland's Findata for the single front door, trusted research environments for access, the European Health Data Space for cross-border legal scaffolding — and the governmental art is sequencing: standards and stewardship funding before platforms, public involvement before controversy, transparency before the tabloid finds the story. Government also carries the failure modes at national scale: care.data shows what a trust collapse costs, and a decade of under-funded registries shows what stewardship neglect costs. The duty is to treat the national data asset as infrastructure — funded, governed, and accounted for like roads — because no market actor will supply its public-good layer (the policy machinery is Chapter 3.2 — Health Policy).

Common failure modes

  • The swamp. Building a lake because storage is cheap and deferring curation forever, until nothing in it can be trusted or found. Fix: no platform without a funded stewardship and catalogue plan; measure use, not terabytes.

  • The bottleneck rebuilt. Centralizing all data work in one team that becomes a queue, or decentralizing into a "mesh" no domain is funded to staff. Fix: choose architecture by where context and capability actually sit, and fund the model you chose.

  • Friction as false safety. Multi-year access processes that ration by stamina while shipping unaudited extracts on request. Fix: single front door, published timelines, enclave-based access with vetted egress.

  • Quality blamed downstream. Cleaning the same defects in every study while the point of creation never hears about them. Fix: feedback loops to data creators, capture paths redesigned, incentives aligned with what audit can verify.

  • Anonymization overclaimed. Marketing de-identified data as risk-free until a re-identification story breaks. Fix: treat de-identification as a spectrum, document residual risk, and match technique to use.

  • The surprise headline. Data uses — especially commercial — that the public first learns about from journalists. Fix: publish the use register, set commercial terms in public in advance, involve patients in access policy.

  • Value unaccounted. A platform that cannot say what its access enabled, defenceless at budget time. Fix: track studies, decisions, and savings enabled by access as part of the access process itself.

  • Hoarding. Institutions guarding data as bargaining power while the combinatorial value of linkage is forfeited system-wide. Fix: make sharing the funded, standardized default and hoarding the position that must justify itself.

Maturity model

Dimension Initiate Develop Standardize Manage Orchestrate
Asset stewardship Datasets accumulate unowned; documentation in departed heads Key assets identified; stewardship ad hoc Inventory with named stewards; curation and metadata funded as standing costs Quality and use of each asset tracked; depreciation acted on Stewardship orchestrated across domains and partners; the asset base managed as a portfolio
Architecture Fragmented silos and shadow extracts Central warehouse or lake serves some needs; queues form Architecture chosen for organizational shape; standards set; domains resourced for their role Bottlenecks measured and rebalanced; platform evolves against analyst-time evidence Federated ecosystem in which domains publish data products to shared standards, nationally linkable
Access Ad hoc requests, personal contacts, shipped extracts Formal process exists but slow and inconsistent Single front door, published timelines, enclave-based access as default End-to-end access time measured and driven down; friction priced as cost Access at research speed across institutions and borders, with auditability strengthening as friction falls
Quality at source Defects found in analysis, fixed downstream forever Quality measured on key datasets Feedback to data creators and capture-path design as standard Quality incentives aligned and audited; improvement tracked at source Quality engineered into the care process; creators see and value what their data becomes
Trust and licence Uses opaque; consent assumed; surprises waiting Legal compliance in place; transparency partial Use register public; patient involvement in access policy; commercial terms set in advance Trust measured; incidents rehearsed; licence actively maintained Public a genuine partner in the data institution; social licence robust to hostile headlines

Checklist

  • Maintain an inventory of significant data assets, each with a named steward and a funded curation plan.
  • Choose warehouse, lake, or mesh by where context, capability, and bottlenecks actually sit — and fund the model chosen.
  • Never commission a data platform without a stewardship, catalogue, and metadata budget.
  • Measure data quality at the point of creation and close the feedback loop to the people who create it.
  • Measure end-to-end access time as a cost, publish targets, and build a single front door for requests.
  • Default to trusted research environments with vetted egress; treat shipped extracts as exceptions and liabilities.
  • Match privacy technique (de-identification, differential privacy, federated approaches) to each use and document residual risk honestly.
  • Publish the register of data uses and involve patients and the public in access policy.
  • Set commercial access terms openly and in advance, including charging and the public system's return.
  • Price access at cost-recovery with public-interest waivers, not monopoly rent and not zero.
  • Track and publish the value access enables — studies, decisions, safety findings, avoided duplication.
  • Apply FAIR principles (findable, accessible, interoperable, reusable) as the stewardship standard for every published dataset.

Key sources

  • OECD — Recommendation on Health Data Governance (2016), the intergovernmental baseline for national health-data frameworks.
  • European Union — European Health Data Space regulation, the emerging legal scaffolding for primary and secondary use across member states.
  • Findata (Finland) — the national health and social data permit authority, the reference implementation of a single front door with secure access.
  • Health Data Research UK — national institute for health-data science, including work on trusted research environments and data access.
  • FAIR principles (Wilkinson et al., 2016) — the findable–accessible–interoperable–reusable standard for data stewardship.
  • The care.data programme and its published post-mortems — the canonical case study in losing the social licence for health-data use.

References

  1. Health data — Wikipedia — https://en.wikipedia.org/wiki/Health_data
  2. Secondary data — Wikipedia — https://en.wikipedia.org/wiki/Secondary_data
  3. Data warehouse — Wikipedia — https://en.wikipedia.org/wiki/Data_warehouse
  4. Data lake — Wikipedia — https://en.wikipedia.org/wiki/Data_lake
  5. Data mesh — Wikipedia — https://en.wikipedia.org/wiki/Data_mesh
  6. Data quality — Wikipedia — https://en.wikipedia.org/wiki/Data_quality
  7. Data governance — Wikipedia — https://en.wikipedia.org/wiki/Data_governance
  8. De-identification — Wikipedia — https://en.wikipedia.org/wiki/De-identification
  9. Differential privacy — Wikipedia — https://en.wikipedia.org/wiki/Differential_privacy
  10. Federated learning — Wikipedia — https://en.wikipedia.org/wiki/Federated_learning
  11. General Data Protection Regulation — Wikipedia — https://en.wikipedia.org/wiki/General_Data_Protection_Regulation
  12. European Health Data Space — Wikipedia — https://en.wikipedia.org/wiki/European_Health_Data_Space
  13. FAIR data — Wikipedia — https://en.wikipedia.org/wiki/FAIR_data
  14. Care.data — Wikipedia — https://en.wikipedia.org/wiki/Care.data
  15. Findata — Finnish Social and Health Data Permit Authority — https://findata.fi/en/
  16. Health Data Research UK — https://www.hdruk.ac.uk/