Chapter 2.6
Evidence Synthesis and Meta-Analysis
Evidence synthesis is the discipline of turning a scattered, uneven, and partly hidden body of studies into a single defensible estimate — because the numbers a decision-maker acts on are almost never the finding of one perfect study, but the pooled verdict of many imperfect ones.
Why this matters in health economics
Every economic evaluation rests on inputs that had to come from somewhere: the relative effect of a treatment, the rate at which a disease progresses, the quality-of-life weight of a health state, the chance of an adverse event. These are the numbers that drive an incremental cost-effectiveness ratio, and where they come from decides whether the ratio is trustworthy. A decision model (see Chapter 2.2 — Modelling) is only as good as the parameters fed into it, and those parameters are the product of synthesis — the deliberate pooling of a whole literature — not of any single trial. This is the chapter that owns "where the numbers in the model come from".
For a director, the stakes are that a synthesis can quietly decide a funding question before the economics is even run. If a review cherry-picks favourable trials, buries a null result, or averages across studies that were never comparable, the effect estimate it produces will be wrong in a specific direction, and every downstream calculation inherits that error with a veneer of rigour. Health technology assessment (HTA) bodies worldwide — England's National Institute for Health and Care Excellence (NICE), Germany's IQWiG, the World Health Organization's (WHO) guideline panels — build their recommendations on syntheses precisely because a single study is too fragile to bear the weight of a national decision.
The stakes are also ethical and about public trust. A well-conducted systematic review is the strongest defence a public body has against being captured by whoever ran the most memorable trial or shouted the loudest. It is a transparent, reproducible method for letting the whole evidence base speak, including the parts that are inconvenient. When a payer can show that its decision followed from a pre-specified, comprehensive, critically-appraised synthesis, it can defend that decision to clinicians, patients, industry, and courts. When it cannot, the decision looks arbitrary — and in health, arbitrary decisions about who gets treated corrode trust fast.
Core concepts
A systematic review is a structured method for locating, appraising, and summarizing all the studies that bear on a specific question, using a pre-registered protocol so the process is reproducible and resistant to bias. It is distinguished from a traditional narrative review by its explicit search strategy, its stated inclusion and exclusion criteria, and its critical appraisal of each included study. The reporting standard most reviews now follow is PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), which asks authors to document every step from search to synthesis so a reader can check what was done.
Meta-analysis is the statistical part of that process: the quantitative pooling of results from several studies into a single combined estimate, usually of an effect size such as a relative risk, odds ratio, or mean difference. Not every systematic review contains a meta-analysis — pooling is only legitimate when the studies are similar enough to combine — but where it is appropriate, it increases precision and can reveal an effect that individual under-powered studies missed. The canonical way to display it is the forest plot: each study a horizontal line showing its estimate and confidence interval, the pooled result a diamond at the foot.
Two statistical models sit behind the pooled estimate. A fixed-effect model assumes every study is estimating one single true effect, and differences between them are only sampling noise. A random-effects model assumes the true effect genuinely varies across studies — because populations, doses, and settings differ — and estimates the average of that distribution, giving more weight to smaller studies and wider, more honest, uncertainty. The choice between them is not cosmetic: it changes the confidence interval and sometimes the conclusion.
The reason the choice matters is heterogeneity — the extent to which study results differ by more than chance would explain. Clinical heterogeneity (different patients or treatments), methodological heterogeneity (different study designs or quality), and statistical heterogeneity (dispersed results) all threaten the meaning of a pooled number. Combining truly heterogeneous studies produces an average that describes no real population — the statistical equivalent of a person with one foot in ice and one in fire being "on average comfortable". Investigating heterogeneity, rather than averaging it away, is often where the real learning lives.
Not all evidence is equal, and the field organizes it into a hierarchy of evidence: systematic reviews of randomized trials at the top, then individual randomized controlled trials, then observational studies, then case series and expert opinion. The hierarchy is a rule of thumb about the risk of bias in a study design, not a guarantee — a poorly-run trial can be worse than a good cohort study. The modern refinement is GRADE (Grading of Recommendations Assessment, Development and Evaluation), which rates the certainty of the evidence for each outcome as high, moderate, low, or very low, and can downgrade even trial evidence for risk of bias, inconsistency, indirectness, imprecision, and publication bias — or upgrade strong observational evidence. GRADE is now used by WHO, Cochrane, and guideline bodies worldwide precisely because it makes the reasoning behind a certainty rating explicit and auditable.
Where head-to-head trials do not exist — B was never trialled against C, only each against placebo — network meta-analysis (also called mixed-treatment or indirect-comparison meta-analysis) combines direct and indirect evidence across a connected network of treatments to estimate all the comparisons, including ones never studied directly. It is indispensable for HTA, where a new drug must be compared with several existing options at once, but it rests on a strong assumption — that the trials are similar enough for indirect comparison to be valid — that must be checked, not assumed.
Two further ideas shape whether a synthesis can be believed. Publication bias is the tendency for studies with positive, significant results to be published, and negative ones to sit in a drawer — so the visible literature systematically overstates effects. It is often screened for with a funnel plot, whose asymmetry can signal missing small negative studies, and it is one reason reviewers search trial registries and grey literature — theses, reports, unpublished data — not just journals. Finally, because evidence keeps arriving, a growing practice is the living review: a synthesis kept continuously up to date as new studies appear, so guidance never drifts years behind the science. The Cochrane collaboration (Cochrane) remains the international reference point for systematic-review methods, and its Handbook is the field's standard manual.
Best practices
Write and register the protocol before you search. A systematic review's credibility comes from deciding the question, the inclusion criteria, the outcomes, and the analysis plan before seeing the results, then registering that protocol publicly (for example in PROSPERO). This is the single strongest guard against the reviewer unconsciously shaping the review to a preferred answer. When you commission a synthesis, insist on a registered protocol and treat its absence as a warning that the conclusions may have been chosen first.
Search comprehensively, including the studies nobody published. A search restricted to a couple of databases and to English-language journals will miss trials and will skew towards positive results. Comprehensive synthesis searches multiple databases, trial registries, conference abstracts, regulatory dossiers, and grey literature, and pursues unpublished data from investigators and manufacturers. The point is not exhaustiveness for its own sake but neutralizing publication bias at the source — the missing studies are disproportionately the negative ones.
Appraise the risk of bias in every included study, and let it change the weight you give the result. Inclusion is not endorsement; each study must be assessed for the specific ways it could mislead — inadequate randomization, unblinded outcome assessment, selective reporting, attrition. A synthesis that pools strong and weak studies with equal weight launders the weak ones' bias into an authoritative-looking average. Use a recognized appraisal tool, report the judgements transparently, and run the analysis with high-risk studies excluded to see whether the conclusion holds.
Decide whether pooling is even legitimate before you compute a pooled number. Meta-analysis is only meaningful when the studies are similar enough in population, intervention, comparator, and outcome to be estimating a comparable effect. If they are not, the honest output is a structured narrative synthesis, not a spurious single number. Ask "are these studies answering the same question?" before asking "what is the combined estimate?" — a precise average of incomparable things is worse than no average.
Choose the fixed-effect or random-effects model deliberately, and justify it. A random-effects model is usually the more honest default in health, where studies genuinely differ, because it produces wider intervals that reflect real-world variation. But it gives small studies disproportionate weight, so where publication bias is suspected it can amplify the very bias you fear. State which model you used and why, and show how the pooled estimate moves under the alternative; a conclusion that flips between the two is not robust.
Investigate heterogeneity rather than averaging it away. When study results diverge more than chance explains, that divergence is information — it may mean the treatment works in one subgroup and not another, or that quality drives the apparent effect. Explore it with pre-specified subgroup analyses and meta-regression, but treat post-hoc subgroups sceptically because they are easy to find and hard to trust. A pooled estimate presented without any examination of why studies disagree is hiding the most useful part of the data.
Screen for publication and reporting bias explicitly. Compare the published literature against trial registries to find studies that were run but never reported, inspect funnel-plot asymmetry where enough studies exist, and consider whether the effect shrinks as study size grows. None of these tests is conclusive, but their absence should lower your confidence in a tidy positive result. The direction of publication bias is nearly always the same — it inflates the apparent benefit — so an unscreened synthesis should be read as an upper bound.
Use network meta-analysis when the decision has many comparators, and check its core assumption. Real funding decisions rarely pit one drug against placebo; they choose among several active options, most of which were never trialled head-to-head. Network meta-analysis fills those gaps, but only if the trials are sufficiently similar for indirect comparison — the "transitivity" assumption — and if direct and indirect evidence agree where both exist. Demand that the analyst show the network, test consistency, and be candid about comparisons resting on thin indirect chains.
Grade the certainty of the evidence, not just the size of the effect. A large effect from low-certainty evidence is a weaker basis for spending public money than a modest effect from high-certainty evidence. GRADE separates these questions by rating certainty per outcome and documenting why it was downgraded or upgraded, turning "the evidence is strong" from an assertion into an audited judgement. Insist that a synthesis feeding a decision carries an explicit certainty rating for each outcome that matters, not a single global verdict.
Report the synthesis so a sceptic could reproduce it. Follow a recognized reporting standard such as PRISMA: give the full search strategy, the flow of studies from identified to included, the excluded studies with reasons, the risk-of-bias judgements, and the data behind every pooled estimate. Transparency is the mechanism by which errors and selective choices are caught before they move a budget. A synthesis whose search cannot be re-run or whose exclusions are unexplained should not anchor a decision, however confident its abstract.
Keep high-stakes syntheses living where the evidence is moving fast. In a fast-changing field — a new drug class, an emerging pathogen, a rapidly-trialled technology — a review is out of date the moment it is published, and stale evidence quietly misdirects money and care. A living review updates continuously against new studies, with pre-agreed rules for when a new trial changes the conclusion. Reserve the effort for decisions large and volatile enough to justify it, and be explicit about the currency date of any synthesis you rely on.
Match the synthesis to the parameter the model actually needs. A decision model needs specific inputs — a relative effect for the base case, a baseline risk for the local population, a utility weight — and each may need its own synthesis with its own inclusion criteria. Pulling a relative effect from randomized trials but a baseline risk from a local registry is legitimate and common; hiding that the two came from different bodies of evidence is not. Align each synthesized input with how it will be used downstream, and cross-reference the evaluation it feeds (see Chapter 2.1 — Economic Evaluation).
Questions to discuss with your team
When we accept an effect estimate as settled, do we know whether it came from the whole literature or only the visible, flattering part of it? This question targets publication bias and selective reporting, the failures that make a synthesis wrong in a predictable direction. The uncomfortable reality is that positive trials are published faster, more often, and more prominently than negative ones, so a review confined to journals will tend to overstate benefit even when every included study is honest. Push the discussion to concrete provenance: were trial registries searched for studies that were run but never reported, was unpublished data sought from manufacturers, was a funnel plot or comparable check done? The tension is that comprehensive searching is expensive and slow, and there is always pressure to accept a clean published estimate and move on. An honest answer does not claim the search was perfect; it states how hard the review looked for the missing studies and treats an unscreened positive result as a likely upper bound rather than the truth.
Are the studies we are pooling actually answering the same question, or have we manufactured a tidy average out of things that do not belong together? This is the heterogeneity question, and it decides whether a pooled number means anything at all. The angles worth surfacing are clinical (different patients, doses, or comparators), methodological (trials mixed with weak observational studies), and statistical (results that disagree by more than chance). The seductive failure is that software will always return a pooled estimate with a confident-looking confidence interval, whether or not the underlying studies are comparable — the number's precision says nothing about its validity. Discuss whether the review investigated why studies disagree or simply averaged them, and whether any subgroup differences were pre-specified or fished for after the fact. An honest answer is willing to conclude that pooling was inappropriate and that a structured narrative synthesis would have been the more truthful output.
How current is this evidence, and what would a new study have to show to change our decision? Syntheses have a shelf life, and the danger is acting on a review whose search closed years ago while the field moved on. Name the currency date explicitly and ask what has been published since, especially for fast-moving technologies where a single large trial can overturn a pooled estimate. The deeper discipline is deciding, in advance, what finding would change your mind — a threshold that turns "keep watching the literature" from a vague intention into an operational trigger for a living update. The tension is between the cost of keeping a synthesis alive and the cost of a decision quietly running on stale evidence; not every question justifies a living review, but every high-stakes, volatile one does. An honest answer states the currency date, names what would reverse the conclusion, and assigns responsibility for noticing when it arrives.
Was the question, and the analysis that would answer it, fixed before anyone saw the results — or could the review have been shaped towards a preferred conclusion? This targets the single strongest guard in the whole discipline: a protocol registered in advance, with the inclusion criteria, outcomes, and pooling plan settled before the data could influence them. The failure it guards against is subtle because it is rarely deliberate — a reviewer who sees the results first can, in good faith, adjust an inclusion boundary, promote a secondary outcome, or switch a model choice in ways that all happen to favour a particular answer. Ask whether a protocol exists in a public registry such as PROSPERO, whether the final review matches it, and whether any deviations were declared and justified rather than quietly made. The tension is that pre-registration is inconvenient — it forbids the "we learned as we went" flexibility that feels like good science — and its absence is easy to excuse. An honest answer can point to the registered protocol, walk through where the review departed from it and why, and treat an unregistered synthesis as one whose conclusions may have been chosen before the evidence was in.
When the decision spans several treatments that were never trialled against each other, do we trust the indirect comparison that stitches them together? Real funding choices usually pit several active options against one another, most never tested head-to-head, so the synthesis leans on a network meta-analysis that borrows strength through shared comparators. The angles to surface are whether the analyst actually showed the network of trials, whether the "transitivity" assumption holds — that the trial populations and background care are similar enough for an indirect comparison to mean anything — and whether direct and indirect evidence agree where both exist. The seductive failure is that the software will return a full set of pairwise estimates however thin or incomparable the underlying chains, lending a fragile indirect result the same authority as a robust direct one. Discuss which contrasts rest on a single weak indirect link and whether those are the very comparisons driving the recommendation. An honest answer distinguishes the well-connected, consistency-tested comparisons from the thin ones, and carries the latter as low certainty rather than burying them in a tidy league table.
Are we treating a big effect and a certain effect as the same thing, when they are not? This separates the size of an estimate from the confidence we are entitled to have in it — the distinction GRADE exists to make explicit. The tension is that a large point estimate is persuasive and easy to spend money on, while the certainty behind it may be low because the evidence is inconsistent, imprecise, indirect, or shadowed by publication bias. Ask whether each outcome that matters carries its own certainty rating, why it was downgraded or upgraded, and whether the recommendation is being driven by the effect size or by an honest reading of how much the evidence can bear. The uncomfortable case is a modest but high-certainty benefit competing against a dramatic but low-certainty one: the discipline is to prefer the evidence you can trust over the number that flatters. An honest answer refuses to collapse "large" and "certain" into a single verdict, and lets the certainty rating, not just the point estimate, govern how much weight the decision puts on the result.
In practice: a health economics example
The National Medicines and Health Technology Committee of Maridia — a fictional lower-middle-income country building out its essential medicines list — must decide which class of blood-pressure-lowering drug to recommend as first-line for uncomplicated hypertension in primary care. Four classes are candidates, all off-patent and affordable, but the clinical question is which prevents the most strokes and heart attacks per patient treated. The committee cannot run its own trials; it must synthesize what the world already knows, then hand a relative-effect estimate to its economics team to combine with local costs and event rates.
The committee's evidence unit begins with a registered protocol: the question, the four comparators, the outcomes (cardiovascular events and all-cause mortality), the inclusion criteria, and the analysis plan, all fixed before searching. The search covers several databases, two trial registries, and WHO regional literature that indexes trials from settings like Maridia's own, because relying only on high-income-country journals would import a population that may not match. The flow is documented PRISMA-style: several thousand records identified, most excluded on title, a few dozen randomized trials included, each appraised for risk of bias. Three large industry-funded trials are found to have reported selectively; the unit flags them and plans a sensitivity analysis without them.
The core difficulty is that almost no trials compared the four classes head-to-head — most tested one drug against placebo or against a single alternative. A simple pairwise meta-analysis can only answer a fraction of the comparisons the committee actually faces. So the unit builds a network meta-analysis, connecting the classes through their shared comparators, and estimates all six pairwise contrasts at once. Before trusting it, they check the network's transitivity — are the trial populations and background care similar enough for indirect comparison? — and test whether direct and indirect evidence agree where both exist. They do, mostly; one contrast rests on a thin indirect chain, which the unit marks as low certainty.
Heterogeneity is real and informative. Trials in older, higher-risk populations show larger absolute benefits than trials in younger cohorts, which matters because Maridia's treated population will skew younger than the trial average. Rather than pool blindly, the unit runs pre-specified subgroup and meta-regression analyses and reports the effect conditioned on baseline risk. A funnel plot on the largest comparison shows mild asymmetry, consistent with a few missing small negative studies; combined with the selective-reporting trials, this leads them to treat the most favourable class's edge cautiously. Under a random-effects model the differences between three of the four classes are small and their confidence intervals overlap heavily.
The unit grades each outcome with GRADE. The evidence that all four classes beat no treatment is high certainty; the evidence that any one class is clearly superior to the others is only low-to-moderate, downgraded for inconsistency, imprecision, and suspected publication bias. That distinction is the decision. The committee concludes that, because the classes are clinically close and all cheap, the economics should be driven by price, tolerability, and supply security rather than by a fragile claim of clinical superiority — a conclusion that only became visible because the synthesis separated effect size from certainty. It hands the economics team a relative-effect estimate with explicit uncertainty, to be turned into a cost-per-event-avoided (see Chapter 2.1 — Economic Evaluation), and commissions a living update: if a large pragmatic trial due in two years shifts a specified contrast beyond a pre-agreed margin, the recommendation is reopened. The single-study causal reasoning inside each trial — whether its own effect estimate is trustworthy — belongs to Chapter 2.3 — Health Econometrics; this committee's job was to pool them well.
Four sector lenses
Startup
A digital-health or device start-up rarely commissions a full systematic review, but it meets synthesis from the other side — as the party whose single promising study must survive an HTA body's synthesis of the wider field. The temptation is to lean on one flattering trial and a narrative review of hand-picked supporting papers, which a serious assessor will see through immediately. The wiser move is to map the existing evidence base honestly and early, so the start-up knows where its product sits in the hierarchy and which comparison a payer will demand. A small firm that presents its evidence in the context of a fair synthesis, and is candid about the certainty of its own data, earns more credibility than one that oversells an isolated result.
Small business
A small but established provider — a group practice, a single specialist clinic, a modest device or diagnostics supplier — has neither the volume of a start-up's investor-backed evidence push nor an enterprise's synthesis team, yet it still has to read syntheses well enough to make steady-state decisions. Its realistic task is consumption, not production: reading the certainty rating on a guideline, checking whether a review's population resembles the patients it actually treats, and noticing when the evidence behind a purchasing or referral habit has quietly gone stale. The trap is deferring wholesale to a vendor's cherry-picked evidence pack or to a single memorable trial, because the practice lacks the time to appraise the wider field. The proportionate move is to lean on the syntheses that trusted bodies — Cochrane, national HTA agencies, professional-society guidelines — already maintain, treat their GRADE certainty ratings as the headline, and ask one disciplined question of any supplier claim: where does this sit against the systematic review, and how current is it?
Enterprise
A large manufacturer, provider group, or insurer runs synthesis as an industrial capability: dedicated evidence-synthesis and health-economics-and-outcomes-research teams producing systematic reviews and network meta-analyses to support submissions across many jurisdictions. The strength is methodological depth and the ability to build a network meta-analysis spanning every relevant comparator; the risk is that the same capability can be turned to producing a review engineered to favour the sponsor's product — through a narrow search, a convenient model choice, or quiet exclusions. Enterprise credibility depends on registering protocols, following PRISMA, submitting reproducible analyses, and inviting the independent replication that HTA bodies increasingly demand. The mature enterprise treats a transparent synthesis that shows genuine uncertainty as more durable than a polished one a reviewer can dismantle.
Government
A ministry, national payer, or HTA body is the principal consumer and referee of syntheses, and its discipline is to demand and to conduct them to a published standard. Bodies such as NICE, IQWiG, and WHO guideline panels build recommendations on systematic reviews and grade the certainty of evidence with GRADE precisely so that decisions are defensible and reproducible rather than captured by the most persuasive submission. Government also carries responsibilities the other sectors do not: to maintain the infrastructure — trial registries, review collaborations, methods guidance — that makes honest synthesis possible, and to keep high-stakes reviews living so national guidance does not drift behind the science. The accountability for what a synthesis leaves out, and for the patients affected by a wrong pooled estimate, ultimately rests here.
Common failure modes
The narrative review dressed as evidence. A hand-picked, unsystematic reading of the literature presented with the authority of a synthesis. Fix: require a registered protocol, an explicit search strategy, and PRISMA-style reporting before trusting any "review".
Pooling incomparable studies. Computing a precise average across trials with different populations, doses, or designs, producing a number that describes no real patient. Fix: assess clinical and methodological heterogeneity first, and use a structured narrative synthesis when pooling is not justified.
Ignoring the studies that were never published. Searching only a couple of databases and only journals, so the missing negative studies inflate the effect. Fix: search registries, regulators, and grey literature, seek unpublished data, and screen for publication bias explicitly.
Averaging away heterogeneity. Reporting a single pooled estimate while suppressing the fact that studies strongly disagree. Fix: investigate divergence with pre-specified subgroup analysis and meta-regression, and report the variation, not just the mean.
Treating inclusion as endorsement. Weighting weak and strong studies equally, laundering bias into an authoritative average. Fix: appraise risk of bias in each study and run the analysis excluding high-risk ones.
Confusing effect size with certainty. Acting on a large effect from low-certainty evidence as if it were settled. Fix: grade certainty per outcome with GRADE and let it, not just the point estimate, govern the decision.
Unchecked indirect comparison. Building a network meta-analysis without testing transitivity or consistency, so an indirect result rests on incomparable trials. Fix: show the network, test the assumptions, and mark thin indirect chains as low certainty.
The stale synthesis. Anchoring a decision to a review whose search closed years ago in a field that has moved on. Fix: state the currency date, check what has appeared since, and keep volatile high-stakes reviews living.
Maturity model
| Dimension | Initiate | Develop | Standardize | Manage | Orchestrate |
|---|---|---|---|---|---|
| Search & selection | Convenience reading of a few known papers | Multiple databases searched, but journals only | Registered protocol, comprehensive multi-source search, PRISMA flow | Registries, regulators, grey literature and unpublished data pursued; searches reproducible | Search infrastructure shared across the portfolio; provenance auditable and re-runnable by third parties |
| Appraisal & pooling | Studies pooled regardless of quality or comparability | Some quality screening; pooling by default | Risk of bias assessed per study with a recognized tool; pooling justified against heterogeneity | Sensitivity analyses by risk of bias; narrative synthesis chosen when pooling is invalid | Appraisal and pooling standards enforced across teams and reconciled with external replication |
| Heterogeneity & bias | Single pooled number, no examination | Heterogeneity noted, not investigated | Subgroup and meta-regression pre-specified; publication bias screened with funnel plots | Heterogeneity explained and used to condition estimates; bias quantified and its direction reported | Heterogeneity insight feeds subgroup-specific decisions system-wide; bias screening standardized across all syntheses |
| Certainty & comparators | Effect size reported as if certain; one comparator | Some quality caveats; pairwise comparisons only | GRADE certainty per outcome; network meta-analysis with assumptions checked | Certainty drives the recommendation; networks validated for consistency and transitivity | Certainty ratings and validated networks coordinated across products, jurisdictions, and partner bodies |
| Currency | Undated review used indefinitely | Currency date recorded | Update schedule agreed for key questions | Living review with pre-agreed triggers for reopening the decision | Living evidence maintained collaboratively across institutions, feeding guidance in near real time |
Checklist
- Confirm a protocol was registered before the search, with pre-specified question, outcomes, and analysis plan.
- Check the search covered multiple databases, trial registries, and grey literature — not journals alone.
- Verify each included study was appraised for risk of bias, and that high-risk studies were tested in a sensitivity analysis.
- Establish that the studies are similar enough to pool; if not, expect a structured narrative synthesis rather than a single number.
- Note whether a fixed-effect or random-effects model was used and why, and whether the conclusion survives the alternative.
- Ask how heterogeneity was investigated, and whether subgroup analyses were pre-specified or post-hoc.
- Look for an explicit publication-bias check (registry comparison, funnel plot) and read a positive result as an upper bound.
- For multi-comparator decisions, confirm any network meta-analysis showed its network and tested transitivity and consistency.
- Require a GRADE (or equivalent) certainty rating for each outcome that matters, separate from the effect size.
- Confirm the synthesis is reported reproducibly (PRISMA), with the search, exclusions, and pooled data available.
- Record the currency date and what new finding would reopen the decision.
- Align each synthesized input with the model and evaluation it feeds, documenting where different inputs came from different evidence.
Key sources
- Cochrane, Cochrane Handbook for Systematic Reviews of Interventions (eds. Higgins, Thomas et al.) — the international standard manual for systematic-review and meta-analysis methods.
- PRISMA Statement (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) — the reporting standard for systematic reviews and meta-analyses — https://www.prisma-statement.org
- GRADE Working Group — the framework for rating certainty of evidence and strength of recommendations — https://www.gradeworkinggroup.org
- NICE Decision Support Unit Technical Support Documents on evidence synthesis and network meta-analysis — methods guidance used in health technology assessment — https://www.sheffield.ac.uk/nice-dsu
- WHO Handbook for Guideline Development — how the World Health Organization uses systematic review and GRADE in global guidelines — https://www.who.int
- Cochrane and the international living-evidence collaborations — reference practice for keeping high-stakes syntheses current.
References
- Systematic review — Wikipedia — https://en.wikipedia.org/wiki/Systematic_review
- Meta-analysis — Wikipedia — https://en.wikipedia.org/wiki/Meta-analysis
- Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) — Wikipedia — https://en.wikipedia.org/wiki/Preferred_Reporting_Items_for_Systematic_Reviews_and_Meta-Analyses
- Forest plot — Wikipedia — https://en.wikipedia.org/wiki/Forest_plot
- Random effects model — Wikipedia — https://en.wikipedia.org/wiki/Random_effects_model
- Heterogeneity (statistics) — Wikipedia — https://en.wikipedia.org/wiki/Heterogeneity_(statistics)
- Hierarchy of evidence — Wikipedia — https://en.wikipedia.org/wiki/Hierarchy_of_evidence
- GRADE approach — Wikipedia — https://en.wikipedia.org/wiki/GRADE_approach
- Network meta-analysis — Wikipedia — https://en.wikipedia.org/wiki/Network_meta-analysis
- Publication bias — Wikipedia — https://en.wikipedia.org/wiki/Publication_bias
- Grey literature — Wikipedia — https://en.wikipedia.org/wiki/Grey_literature
- Cochrane (organisation) — Wikipedia — https://en.wikipedia.org/wiki/Cochrane_(organisation)
- Cochrane Handbook for Systematic Reviews of Interventions — Cochrane — https://training.cochrane.org/handbook
- PRISMA Statement — PRISMA — https://www.prisma-statement.org
- GRADE Working Group — https://www.gradeworkinggroup.org