Critical Appraisal: How to Judge Whether a Study's Results Can Be Trusted
Critical appraisal is the structured process of reading a research report and judging how much confidence its methods can support. It answers three separate questions: are the results likely to be biased, what do the results actually say, and do they apply to the question you are asking? It is not a verdict on whether the topic is interesting or whether you agree with the conclusion.
The most consequential choice is matching the tool to the study design. RoB 2 is the reference standard for randomised trials, ROBINS-I for non-randomised studies of interventions, QUADAS-2 for diagnostic test accuracy, PROBAST and QUIPS for prediction and prognosis, and AMSTAR 2 or ROBIS for published systematic reviews. The CASP and JBI checklists cover a wide range of designs in plainer language. Reporting guidelines such as CONSORT, STROBE and PRISMA are not appraisal tools; they tell authors what to include in a manuscript.
Risk of bias is a property of a result, not of a paper, so appraise each outcome you plan to use rather than rating the study once. Record the specific sentence, table or registry field behind every judgment, mark a domain as "no information" where the paper is silent, have two people appraise independently and reconcile, and carry the judgments into your synthesis. The sections below cover the checklists, the step-by-step method, worked examples, and where to get training and support.
Checklists and Tools for Appraising Different Study Designs
Critical appraisal is the structured process of reading a research report and judging how much confidence its methods can support. You are not asking whether you agree with the conclusion or whether the topic is interesting; you are asking whether the design, conduct, analysis, and reporting were sound enough that the reported findings are likely to reflect what actually happened, and whether the study population and setting resemble the one you care about. In practice it means separating three questions that readers often blur together: Are the results likely to be biased? What do the results actually say, including how precise they are? And do they apply to my question? Everything else in appraisal is machinery for answering those three questions consistently.
You will see the same activity described under several names, and they are close to interchangeable in everyday use: quality assessment, methodological quality appraisal, risk of bias assessment, and, in teaching contexts, "reading a paper critically." The terms are not perfectly synonymous, though, and the distinction matters when you write up a review. "Quality" tends to describe how well a study was conducted and reported against a standard of good practice; "risk of bias" is narrower and asks specifically whether the study's design or execution could have pushed the estimated effect away from the truth. Modern review methods have moved toward risk of bias language because a study can be reported beautifully and still be at high risk of bias, or reported poorly while having been conducted well. Related but separate concepts include internal validity (were the results correct for the people studied?), external validity or applicability (do they transfer?), and certainty of evidence, which is a judgment about a whole body of studies rather than one paper.
The most consequential choice you make is matching the tool to the study design, because a checklist written for randomized trials asks questions that a cohort study cannot answer and skips the questions that decide whether a cohort study is credible. For randomized controlled trials, the Cochrane risk of bias tool, currently RoB 2, is the reference standard in systematic reviews and works through domains covering the randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result. For non-randomized studies of interventions, ROBINS-I is built around the idea of an emulated target trial and puts confounding and selection into the study at the front. Observational studies appraised outside an intervention framing are often handled with the Newcastle-Ottawa Scale for cohort and case-control designs, or AXIS for cross-sectional studies. For diagnostic test accuracy, QUADAS-2 covers patient selection, the index test, the reference standard, and flow and timing. Prognostic factor studies have QUIPS, and prediction model studies have PROBAST. Systematic reviews themselves are appraised with AMSTAR 2 or ROBIS. Qualitative research has the CASP qualitative checklist and the JBI qualitative tools, with COREQ used more for reporting completeness; mixed-methods studies have the MMAT. Economic evaluations are commonly appraised with the Drummond or CHEC criteria, with CHEERS as the reporting standard. Case reports have CARE, and prevalence studies have a dedicated JBI tool. Two general-purpose families deserve mention because they are what most students and clinicians actually reach for: the CASP checklists, which exist in a version for most major designs and are written in plain language, and the JBI critical appraisal tools, which cover an unusually wide range of designs and are the basis for JBI-style reviews.
A word on reporting guidelines, which are frequently confused with appraisal tools. CONSORT, STROBE, PRISMA, and their relatives tell authors what to include in a manuscript. They are not appraisal instruments, and scoring a paper against CONSORT tells you about the write-up rather than the study. Use them to identify what information is missing, then decide within your appraisal tool how to handle that missing information, usually as "no information" rather than as a judgment of high or low risk.
The mechanics of doing an appraisal are more mundane than the literature suggests. Read the full text once for comprehension before touching a checklist, then extract the study's design, population, intervention or exposure, comparator, outcomes, and analysis into a short structured summary so that you are appraising the study rather than your memory of it. Choose the tool that matches the design and read its guidance document, not just the item list, because most tools define their terms narrowly. Work item by item, and for each judgment record the specific sentence, table, or figure that justifies it; a judgment without a supporting quotation cannot be checked by a co-reviewer or a reader. Where the paper is silent, say so explicitly instead of assuming the worst or the best. Then look at your item-level judgments together and reach an overall judgment for the outcome you care about, remembering that risk of bias is outcome-specific: blinding may be irrelevant for mortality and decisive for a patient-reported symptom score. In review work, two people appraise each study independently, compare, resolve disagreements by discussion, and involve a third person when discussion fails. Dedicated review software helps here mainly by keeping the two appraisals from contaminating each other, storing the supporting quotations next to each judgment, flagging conflicts for reconciliation, and generating the traffic-light and summary figures directly from the recorded data so the published figure and the underlying judgments cannot drift apart.
A short illustrative example makes the pattern concrete. Suppose you are appraising a hypothetical randomized trial of a group exercise program for older adults using RoB 2, with mobility measured by a performance test as your outcome of interest. Under the randomization domain, you find the paper reports computer-generated sequence and sealed opaque envelopes, and the baseline table shows no unexpected imbalance, so you record a low-risk judgment with those quotations attached. Under deviations from intended interventions, participants clearly knew which group they were in, which is unavoidable in an exercise trial, so your judgment turns on whether the analysis was intention-to-treat and whether co-interventions differed; if the paper reports an intention-to-treat analysis and no unusual co-interventions, some concerns is a defensible judgment rather than high risk. Under missing outcome data, you compare the number randomized with the number analyzed and check whether attrition differed between arms and whether reasons were reported. Under outcome measurement, the key question is whether the assessors performing the mobility test were blind to allocation, and the answer often is not stated, so you record "no information" and explain the implication. Under selection of the reported result, you look for a prospectively registered protocol and compare its outcome list and time points against what was published. The overall judgment follows from the weakest domain, and your write-up says which domain drove it. Notice that none of this required an opinion about whether exercise programs work; the appraisal is about the study, not the intervention.
Critical appraisal topics, or CATs, are the short standardized write-ups used in teaching and in clinical practice to answer one focused question. A CAT typically states the clinical scenario, the question in PICO form, the search that was run, the single best study or two identified, the appraisal of that study, the numerical results with confidence intervals, and a bottom-line statement with a date and an expiry. Typical topics are narrow and answerable: whether one imaging test outperforms another for a specific presentation, whether a particular screening interval is supported for a defined age group, whether a rehabilitation protocol changes a functional outcome after a specific surgery, whether a point-of-care test is accurate enough to rule out a condition in primary care, or whether a risk prediction model has been validated in a population resembling yours. Weak topics are broad ones, since "is nutrition important in recovery" cannot be answered by appraising a paper.
Formatting a written appraisal follows a stable pattern regardless of venue. Open with the full citation and a one- or two-sentence statement of the question the study addressed and the question you brought to it. Follow with a brief, neutral summary of the design, participants, comparison, outcomes, and headline results, in your own words and without evaluation, so a reader knows what you are appraising. Then move into the appraisal proper, organized by the domains of your chosen tool, with a heading or short paragraph per domain, each stating your judgment and the evidence for it. Next, address applicability: how the study population, setting, comparator, and outcome definitions compare to your context. Close with a bottom line that states what you now believe, how confident you are, and what would change your mind. Attach the completed checklist as an appendix or table rather than reproducing every item in the prose.
A critical appraisal essay is the same content in continuous prose, and the failure mode to avoid is marching through checklist items as a list of yes-or-no verdicts. Use the checklist as scaffolding during your reading, then write argumentatively: identify the two or three methodological features that actually determine how much weight the study deserves, and spend your words on those, dealing with the rest briefly. For each concern, name the flaw, explain the mechanism by which it could distort the result, and state the likely direction and rough magnitude of the distortion where you can reason about it. Distinguish problems that could plausibly overturn the finding from problems that are merely untidy reporting. Quote sparingly and specifically. Be as explicit about the study's strengths as its weaknesses, because an appraisal that finds only faults reads as a hunt rather than a judgment. And keep your conclusion proportionate to your evidence: the honest end point of many appraisals is that the study is informative but not decisive, and saying so precisely is a better outcome than manufacturing a verdict.
How Do the Main Appraisal Checklists Compare?
There is no single all-purpose appraisal instrument. The main options differ along three axes: the study designs they cover, whether they produce structured judgments or a numeric score, and whether they appraise individual studies or a whole synthesis. In practice, the choice is driven almost entirely by what kind of study you are holding. A randomized trial, a cohort study, a diagnostic accuracy study, a prognostic model, a qualitative study, and a published systematic review each have characteristic ways of going wrong, and each has a tool built around those specific failure modes. Applying a trial-oriented tool to a cohort study, or vice versa, produces judgments that look rigorous but interrogate the wrong things.
The second axis matters more than newcomers expect. Domain-based tools ask you to make a reasoned judgment in each of several prespecified domains, randomization, missing data, outcome measurement, and so on, and to record the information that supports each judgment. Scale-based tools award points and total them. Numeric totals are convenient for sorting a table but they implicitly treat every item as equally important, and a study can lose points on reporting details while remaining sound on the domains that actually threaten its result. Current methodological guidance for reviews of interventions favors domain-based judgments over summary scores, which is why older scales that reduce a trial to a single figure have largely fallen out of use in formal synthesis, even though they still appear in narrative reviews.
| Instrument | Designs it is built for | Output format | Typical application |
|---|---|---|---|
| Cochrane RoB 2 | Randomized trials, including cluster and crossover variants | Domain-level judgments (low, some concerns, high) plus an overall judgment, driven by signaling questions | Intervention reviews where the included evidence is predominantly randomized; supports per-outcome appraisal |
| ROBINS-I | Non-randomized studies of interventions, including cohort and self-controlled designs | Domain-level judgments against a hypothetical target trial, including confounding and selection into the study | Reviews that include observational intervention evidence and need a common framework with trials |
| Newcastle-Ottawa Scale | Cohort and case-control studies | Star-based points across selection, comparability, and outcome or exposure ascertainment | Etiologic and prognostic observational reviews; still common in epidemiology journals despite its scoring approach |
| JBI critical appraisal checklists | A family of design-specific checklists spanning trials, cohorts, case series, prevalence, qualitative, and economic studies | Item-level yes/no/unclear/not applicable responses with an include, exclude, or seek-further-info decision | Mixed-design reviews needing one consistent format across many study types |
| CASP checklists | Design-specific checklists for trials, cohorts, case-control, diagnostic, qualitative, and reviews | Short question sets answered in prose; no scoring | Teaching, journal clubs, and rapid reading of single papers rather than formal synthesis |
| QUADAS-2 | Diagnostic test accuracy studies | Domain-level risk of bias plus separate concerns about applicability | Test accuracy reviews where index test, reference standard, and patient flow drive the result |
| AMSTAR 2 and ROBIS | Published systematic reviews treated as included studies | AMSTAR 2 flags critical and non-critical weaknesses; ROBIS gives phase-based domain judgments | Overviews of reviews and umbrella reviews |
| PROBAST and QUIPS | Prediction model studies (PROBAST) and prognostic factor studies (QUIPS) | Domain-level judgments covering participants, predictors, outcome, and analysis | Prognosis reviews, where confounding and analysis choices differ from intervention questions |
The third axis is what the appraisal feeds into. Appraising individual studies is not the same as rating the certainty of a body of evidence. Frameworks such as GRADE take study-level risk of bias as one input alongside inconsistency, indirectness, imprecision, and publication bias. If your protocol commits to a certainty rating, choose an appraisal tool whose output can be carried forward into that judgment, domain-based tools with per-outcome judgments make this straightforward, while a single composite score generally does not.
How you decide between two defensible options usually comes down to fit with your evidence base and your reviewers. Consider whether the tool covers every design you expect to include, or whether you will need a second instrument and a plan for reporting both. Consider whether it appraises per outcome, which matters when a study is sound for its primary endpoint but not for a secondary one you are extracting. Consider training burden: signaling-question tools are more prescriptive and more reproducible once learned, but they take longer per study and reward a pilot round on a handful of papers before the full run. Simpler checklists move faster and are easier to hand to a new reviewer, at the cost of more variation between appraisers.
Whichever instrument you pick, the mechanics are the same. Name the tool and version in the protocol before appraisal starts. Have two reviewers appraise independently, record the text or table that justifies each judgment rather than the judgment alone, and document how disagreements are resolved. Appraisal is not a filter to be applied silently: report the per-study results in full and state whether they changed the analysis, whether through subgroup or sensitivity analyses restricted to studies at lower risk of bias, or through the certainty rating. Review software helps here mainly by keeping the two appraisers' forms separate until reconciliation, storing supporting quotes with each domain, and exporting the completed judgments as summary and traffic-light figures for the manuscript.
Using Critical Appraisal in Systematic Reviews and Meta-Analyses
Critical appraisal is the step where you judge how much confidence the methods of each included study can support, independent of what the study found. In a systematic review it is not an optional add-on at the end; it belongs in your protocol, it runs alongside data extraction, and its output feeds directly into how you synthesize and how you word your conclusions. Registering your appraisal plan in advance, in PROSPERO or your protocol document, also protects you against the temptation to appraise more harshly the studies whose results you dislike.
Choose the instrument by study design, not by convenience. Appraisal tools are design-specific, and using the wrong one produces judgments that reviewers and journal editors will discount. For randomized trials, the Cochrane risk-of-bias tool (RoB 2) is the current standard, organized by domain: randomization process, deviations from intended interventions, missing outcome data, outcome measurement, and selection of the reported result. For non-randomized studies of interventions, ROBINS-I works through pre-intervention, at-intervention, and post-intervention domains, including confounding and selection into the study. Diagnostic accuracy studies call for QUADAS-2; prevalence and other observational designs are often handled with the JBI checklists; prognostic model studies with PROBAST. If you are appraising other systematic reviews rather than primary studies, AMSTAR 2 or ROBIS are the relevant instruments. Where your review includes several designs, apply the matching tool to each and report the results separately rather than forcing everything into one scale.
Appraise at the outcome level, not just the study level. This is the single most common technical error in otherwise competent reviews. A trial can be at low risk of bias for mortality, which is hard to misclassify and rarely missing, and at high risk for a patient-reported symptom score assessed by unblinded participants with substantial dropout. RoB 2 and ROBINS-I are both designed to be applied to a specific result, one outcome, one comparison, one time point. Practically, this means your appraisal grid has a row per study-outcome pair, not per study. If you plan three meta-analyses, expect to make three sets of judgments for any study that contributes to all three.
Use two independent appraisers and record the disagreements. The standard method is for two reviewers to appraise each study independently, then reconcile. Reconciliation is a discussion aimed at finding out why you differed, usually one of you saw a methods detail the other missed, or you interpreted the signaling question differently, with a third reviewer or the senior author arbitrating anything that stays unresolved. Before you start, calibrate: appraise two or three included studies together as a group, argue about the borderline domains, and write down the decision rules you settle on. Many reviews also report a chance-corrected agreement statistic such as Cohen's kappa; if you plan to, decide up front, because it has to be calculated on the pre-reconciliation judgments.
Justify every judgment in free text. A domain rating with no supporting quote is not auditable. For each domain, record the specific sentence, table, or absence of information that drove your rating, along with the page or section. Distinguish clearly between "the study did this badly" and "the study did not report it", those are different situations, and the second may be resolvable by contacting authors or checking the trial registration and protocol for the pre-specified outcomes. Registry checks are also how you populate the selective-reporting domain, which cannot be assessed from the published paper alone. Your review software should let you attach these notes and source locations to each judgment so they carry through to the final report without retyping.
Feed appraisal into synthesis rather than into inclusion. Excluding studies because they scored poorly is generally discouraged; it discards information and the threshold is arbitrary. The recommended approaches are to stratify or to test sensitivity. Present the pooled estimate for all studies, then present it again restricted to studies at low risk of bias, and see whether the direction and magnitude hold. You can also enter risk of bias as a subgroup variable or a meta-regression covariate, though with few studies these analyses have little power and should be described as exploratory. If nearly all your studies sit at high risk of bias in the same domain, a pooled estimate may not be defensible at all, and a structured narrative synthesis is the more honest output. Avoid numeric quality scores that sum across domains: they weight unequal problems equally and can mask a fatal flaw behind an acceptable total.
Report it so a reader can re-derive your conclusions. PRISMA 2020 asks you to state which tool you used, how many appraisers worked independently, how conflicts were resolved, whether any automation tools assisted, and how the results were used in synthesis. Present per-study, per-domain judgments in a table or traffic-light figure, with the domain-level summary alongside the overall rating, and put the supporting justifications in a supplementary file. Finally, connect appraisal to certainty: in GRADE, risk of bias is one of the five domains that can lower your confidence in a body of evidence, so your appraisal results should visibly explain any downgrade in your summary-of-findings table. If your conclusions section describes evidence as limited, the reader should be able to trace that word back to specific domain judgments.
What Is Critical Appraisal in Research?
Critical appraisal is the structured process of reading a study and judging how much confidence its design and conduct warrant. It sits between screening and synthesis in a systematic review: screening decides whether a study belongs in your review at all, data extraction records what the study reported, and critical appraisal asks whether the reported result is likely to reflect a real effect or whether some feature of the study's methods could have pushed the answer in a particular direction. It is not a verdict on whether a paper is worth reading, and it is not a judgment about the authors. It is an assessment of specific, named threats to validity in a specific study, for a specific question.
The core distinction to hold onto is between risk of bias and reporting quality. Risk of bias concerns features that could systematically distort the effect estimate, how participants were assigned to groups, whether outcome assessors knew which group they were assessing, whether everyone randomized was accounted for, whether the analysis followed a pre-specified plan. Reporting quality concerns whether the paper tells you enough to make those judgments. The two often travel together, and thin reporting frequently forces an "unclear" or "no information" rating rather than a low-risk one, but they are different problems. A well-conducted trial can be badly written up, and a beautifully written paper can describe a badly designed study. Modern appraisal tools generally ask you to judge the study, not the manuscript, which is why protocols, trial registry records, and supplements are legitimate sources when the main article is silent.
Appraisal is design-specific, and this is where most reviews go wrong. The threats to a randomized trial are not the threats to a cohort study, and neither resembles the threats to a diagnostic accuracy study or a qualitative study. Randomized trials are usually appraised for sequence generation, allocation concealment, blinding, missing outcome data, outcome measurement, and selective reporting. Non-randomized studies of interventions add confounding and selection into the cohort as central concerns, along with how exposure and outcome were classified and whether the analysis handled time-varying factors sensibly. Diagnostic accuracy studies turn on patient selection, the index test, the reference standard, and the flow and timing between them. Prevalence, prognostic, economic, and qualitative studies each have their own frameworks. In practice this means you should decide during protocol writing which instrument applies to which study design in your review, and expect to run more than one instrument if your inclusion criteria admit mixed designs.
The instruments themselves fall into a few recognizable families. Domain-based tools, the Cochrane risk-of-bias tools for trials, ROBINS-I for non-randomized intervention studies, QUADAS-2 for diagnostic accuracy, ask signaling questions within each domain and lead you to a judgment per domain, then to an overall judgment. Checklist tools, including the JBI critical appraisal checklists and the CASP checklists, walk through items and are often used in teaching and in reviews with wide design heterogeneity. Scale-based tools that sum items into a single score have fallen out of favor for interventions, because summing unrelated items weights them arbitrarily and can hide a fatal flaw inside a respectable total. AMSTAR 2 and ROBIS occupy a separate niche: they appraise systematic reviews themselves, which matters if you are writing an overview of reviews. Whichever family you use, name the tool, the version, and any adaptations in your methods section, and archive the exact form your team completed.
Procedurally, appraisal is done at the study level for each outcome that matters, not once per paper. A trial can be at low risk of bias for mortality and at high risk for a subjective symptom score assessed by unblinded participants, and collapsing that into a single per-study rating loses the information your synthesis needs. Two reviewers should appraise independently and then reconcile, with a third person or a pre-agreed rule for deadlocks. Pilot the tool on three to five included studies before the full round; this is where you discover that your team reads "unclear" differently or that a signaling question doesn't map onto your field's conventions. Write down the decision rules you settle on and keep the free-text justification for every non-obvious judgment, because reviewers and readers will ask why a particular study was downgraded, and "we remembered discussing it" is not an answer you can defend a year later.
Finally, appraisal earns its keep only if it changes something downstream. The judgments should feed into how you present results, typically a risk-of-bias figure or table, a narrative account of where the weaknesses cluster, and sensitivity analyses that show whether conclusions hold when higher-risk studies are removed. They also feed into certainty-of-evidence assessment, where study limitations are one consideration alongside inconsistency, indirectness, imprecision, and publication bias. Two habits to avoid: excluding studies on appraisal grounds after they have already met your inclusion criteria, unless your protocol said so in advance and gave a threshold; and appraising every study, then never mentioning the ratings again. If your appraisal has no visible consequence in the results or discussion, readers cannot tell whether it informed your conclusions or was performed to satisfy a reporting checklist.
How to Appraise a Research Paper Step by Step
Critical appraisal is the point in a review where you stop asking what did this study find and start asking how much should this finding move me. The sequence below works whether you are appraising a single paper for a journal club or working through a hundred records for a systematic review. The core discipline is the same: separate what the study did from what the authors concluded, judge the methods against the question you are asking, and leave behind a record that someone else could follow to the same judgment.
Step 1: Assemble the full record before you judge anything. The published article is rarely the whole study. Pull the supplementary appendices, the trial registry or protocol entry, any statistical analysis plan, and prior publications from the same cohort or trial. For older or sparsely reported papers, check whether a conference abstract, dissertation, or regulatory document covers the same data. Appraising from the PDF alone means you are appraising the report, not the study, and you will systematically mistake poor reporting for poor conduct, or miss discrepancies that only appear when the registry sits next to the results section.
Step 2: Identify the design as executed, not as labeled. Papers routinely describe themselves inaccurately. A study called a "randomized trial" may have allocated by alternate assignment; a "cohort study" may be a retrospective chart review with an assembled comparison group; a "systematic review" may have searched one database without a protocol. Read the methods and reconstruct what actually happened: how participants entered the study, how groups were formed, when outcomes were measured relative to exposure, and who controlled the assignment. Everything downstream depends on getting this right, because the design determines which appraisal instrument applies.
Step 3: Choose the instrument that matches the design and the question. Use a randomized-trial tool for randomized trials, a nonrandomized-intervention tool for cohort and other non-randomized comparative designs, a diagnostic-accuracy tool for test accuracy studies, and a review-appraisal tool for existing syntheses. Structured domain-based instruments such as RoB 2, ROBINS-I, QUADAS-2, AMSTAR 2, and the JBI and Newcastle-Ottawa suites each embed assumptions about what can go wrong in that design; using the wrong one produces judgments about the wrong threats. Do not mix instruments within a design category in the same review unless you can justify it, and record the version you used, since these tools are revised.
Step 4: Fix the target of your appraisal before you begin. Risk of bias is a property of a result, not of a paper. A trial can be at low risk of bias for a hard registry-linked outcome and high risk for a self-reported secondary outcome measured by unblinded assessors. Decide which outcome, at which time point, and under which effect of interest, you are appraising, and appraise that. If your review has several outcomes carried into synthesis, you will need a separate appraisal for each. This is the single most common procedural error in appraisal, and it is also the reason two appraisers who read the same paper carefully can disagree completely without either being wrong.
Step 5: Read the methods and results before the abstract and discussion. Read in the order that protects you from the authors' framing. Methods first, then results tables, then figures, then the abstract, then the discussion last. Note the sample size actually analyzed against the number enrolled, whether the primary analysis was intention-to-treat or per-protocol, how missing data were handled, and whether the outcome definition in the results matches the one in the methods. Check that the numbers in the abstract can be found in the tables. Discrepancies here are informative in themselves.
Step 6: Work through each domain and write down the evidence for your judgment. For every domain, quote or cite the specific sentence, table, or registry field you relied on, then record the judgment. "Low risk of bias: allocation by central web-based randomization, p. 4, col. 2" is a usable record; "Low" alone is not. Where the paper says nothing, judge it as no information rather than as adequate. Silence about blinding is not evidence of blinding, and the distinction between "not done" and "not reported" matters when you later decide whether to contact authors or to run a sensitivity analysis.
Step 7: Compare what was planned with what was reported. Put the registry entry or protocol beside the paper and check the primary outcome, the time points, the analysis population, and the planned subgroups. Look for outcomes that appear in the paper but not the protocol, outcomes that were pre-registered but never reported, changes in primary outcome after enrollment began, and subgroup analyses presented as if planned. Note the registration date relative to the first enrollment date. This comparison is the only reliable way to assess selective reporting, and it cannot be done from the article text alone.
Step 8: Judge direction and likely magnitude, not just presence. A concern only matters if it could plausibly change your interpretation. Ask which way each identified problem would push the estimate and how far it would have to push to alter your conclusion. Unblinded assessment of a subjective outcome typically favors the treated group; differential loss to follow-up may push either way depending on who left; confounding by indication in observational treatment comparisons often runs against the treatment. Writing this reasoning down is what turns an appraisal from a checklist score into something useful for interpretation.
Step 9: Appraise in duplicate and reconcile explicitly. Two independent appraisers, then discussion, then a named third party for unresolved disagreements. Before you start at scale, pilot the tool on a handful of included studies together, compare judgments, and write decision rules for the recurring ambiguities specific to your topic, such as what counts as adequate adjustment for a given confounder or how much attrition triggers a concern. Calibration up front does more for consistency than adjudicating hundreds of disagreements later. Record reconciled judgments separately from the initial independent ones.
Step 10: Carry the appraisal into the synthesis, or it was wasted effort. Appraisal that ends in a color-coded figure and is never mentioned again has no bearing on your conclusions. Plan in advance how judgments will be used: stratifying or restricting meta-analyses by risk of bias, running sensitivity analyses that exclude studies with serious concerns, informing the study-limitations judgment in a certainty-of-evidence assessment, or in narrative reviews, giving greater interpretive weight to better-conducted studies and saying so. Note also that appraisal of individual studies and assessment of certainty across a body of evidence are different tasks; the first feeds the second but does not replace it.
How long this takes varies enormously with design, reporting quality, and appraiser experience, and any general figure would be misleading. A well-reported trial with an accessible protocol and a hard outcome can be appraised quickly; a sparsely reported nonrandomized study requiring confounder-by-confounder judgment and author correspondence can take many times longer. Budget by expected mix of designs and reporting quality in your included set, not by paper count, and expect the first several appraisals in any new review to be slower while your decision rules settle.
Training Courses and Further Reading on Appraising Evidence
Critical appraisal is a skill built by doing it repeatedly with feedback, not by reading about it once. That said, structured training shortens the learning curve considerably, and the right starting point depends on what you are appraising and why. A clinician who needs to judge whether a single trial should change practice needs something different from a review author who must apply a formal risk-of-bias instrument to forty studies and defend every judgment to peer reviewers. Before signing up for anything, get clear on which of those you are: it will tell you whether you need a short introduction to reading papers, or in-depth training in one specific tool.
For a low-commitment start, the Critical Appraisal Skills Programme (CASP) publishes free checklists for randomized trials, systematic reviews, cohort and case-control studies, qualitative research, diagnostic accuracy studies, and economic evaluations. The checklists are widely used in teaching because they are short and prompt the reader through the main questions in plain language. CASP also runs paid workshops. The Centre for Evidence-Based Medicine at Oxford publishes free teaching materials, worksheets, and the Catalogue of Bias, which is a useful reference when you can name that something is wrong with a study but cannot articulate which bias it is. Students 4 Best Evidence, aimed at students and early-career researchers, hosts short explanatory posts on appraisal concepts and is a reasonable place to fill in gaps.
For more structured training, Cochrane Training offers both free resources and Cochrane Interactive Learning, a modular paid course covering the systematic review process including risk-of-bias assessment and GRADE. JBI runs training programs in its own approach to evidence synthesis, including qualitative and mixed-methods appraisal, through its Comprehensive Systematic Review Training Program and endorsed training centers. The GRADE Working Group maintains guidance and training materials for rating certainty of evidence, and GRADE workshops run periodically through affiliated centers. In the United States, the AHRQ Evidence-based Practice Center Methods Guide for Comparative Effectiveness Reviews is free and specifies how appraisal is expected to be done in AHRQ-funded reviews, which makes it worth reading if your work will be judged against those standards. The Campbell Collaboration serves the equivalent function for social, behavioral, educational, and criminal justice research, where the study designs and the relevant biases differ from clinical medicine.
Course fees, formats, and time commitments vary substantially between providers and change over time, so check current details directly rather than relying on secondhand figures. Before paying, check whether your institution already holds a license: many academic medical centers and universities have institutional access to Cochrane Interactive Learning or run their own internal appraisal workshops through the library. Your medical librarian or research support office is the fastest way to find out, and librarians are often trained appraisers themselves who will sit down with a paper and work through it with you.
If you prefer books, a few are standard. Trisha Greenhalgh's How to Read a Paper is the usual first recommendation for clinicians because it is organized by study type and assumes no prior methods training. Users' Guides to the Medical Literature, edited by Guyatt and colleagues, is denser and more comprehensive, and is the reference to reach for when you need to understand why a particular question is on a checklist. Straus and colleagues' Evidence-Based Medicine: How to Practice and Teach It is oriented toward people who have to train others. Testing Treatments, freely available online, is written for a general audience and is useful for explaining appraisal concepts to patients, students, or committee members without a research background.
For tool-specific work, go to the source documentation rather than a secondary summary. The Cochrane Handbook for Systematic Reviews of Interventions covers RoB 2 and ROBINS-I, and the tool developers publish separate detailed guidance documents with worked examples and signaling questions. The JBI Manual for Evidence Synthesis covers the JBI appraisal instruments. AMSTAR 2 and ROBIS have their own published guidance for appraising systematic reviews, and QUADAS-2 for diagnostic accuracy studies. One point of frequent confusion: reporting guidelines such as PRISMA, CONSORT, and STROBE, hosted by the EQUATOR Network, are not appraisal tools. They describe what a paper should report, not whether the study was well conducted, and using them as quality scores will produce misleading judgments. Reading them is still worthwhile, because knowing what should have been reported makes gaps easier to spot.
The practice component matters more than the reading. Appraise papers in pairs and compare judgments before discussing them, which surfaces where your interpretations of a domain diverge. Run calibration exercises on two or three papers before starting a review so the team agrees on how borderline cases are handled, and write down the decisions you reach. A journal club with an explicit appraisal focus, where someone is assigned to present the methods rather than the conclusions, gives you regular reps. Keep a note of judgments you later revised and why. To stay current, watch for updates to the tools themselves, since instruments are revised and older versions remain in circulation long after newer ones are published.
Where to Get Extra Support With Appraising Studies
Critical appraisal is a judgment task, not a data-entry task, and even experienced reviewers hit studies they cannot confidently rate. Asking for help is normal practice on a well-run review, not a sign that you are out of your depth. The useful question is which kind of help you need: someone who knows the appraisal instrument, someone who knows the clinical or subject content, someone who knows the statistics, or simply a second pair of eyes to test whether your reasoning holds up. Naming the gap first will save you from sending a vague question to the wrong person.
For most teams the first stop is a research librarian or information specialist. In the United States, academic medical centers, university libraries, and many hospital systems employ librarians who support systematic reviews as a core service, and appraisal support is often part of that remit. They can help you identify which instrument matches your included study designs, track down the full guidance document behind a checklist, find worked examples, and point you to the version of a tool that journals in your field currently expect. If your institution has a formal systematic review service, ask early what it covers; some will join the review team as collaborators, others will advise on methods without appraising studies themselves.
Read the developer's own guidance before you look for outside help, because a surprising share of appraisal problems are answered there. Instruments such as the Cochrane risk of bias tools, ROBINS-I, QUADAS-2, the JBI critical appraisal checklists, AMSTAR 2, and the GRADE approach are all accompanied by manuals or handbooks that define each signaling question, explain what counts as adequate reporting, and give examples of how to reach a domain-level judgment. A checklist stripped of its guidance document invites inconsistency across your team. Many tool websites also list a contact address or maintain a frequently asked questions page, and the Cochrane Handbook, the JBI Manual for Evidence Synthesis, and the GRADE handbook are freely available online and worth keeping open while you appraise.
Formal training is worth the time if you expect to run more than one review. Options include Cochrane's online learning materials and workshops, JBI training programs delivered through its collaborating centers, GRADE workshops, short courses run by schools of public health and epidemiology, and courses offered through professional societies such as the Medical Library Association, the Society for Research Synthesis Methodology, and the Campbell Collaboration for social and behavioral topics. Cost, length, and prerequisites vary widely, so check what a course actually covers before enrolling; some teach a single instrument in depth, others survey the whole review process and touch appraisal only briefly. A cheap and effective supplement is to appraise two or three already-published studies independently with a colleague, then compare and argue through every disagreement before you start on your real included studies.
Methodological consultation is a separate resource from training. Institutions with a Clinical and Translational Science Award hub, a biostatistics core, or a research design consultation service can usually give you an appointment with a statistician or methodologist. Bring specific questions: whether a cluster-randomized trial handled the unit of analysis appropriately, whether a particular approach to missing data leaves a study at risk of bias, whether the confounders adjusted for in an observational study are the ones your review cares about, or whether a crossover or stepped-wedge design raises issues your instrument does not obviously cover. Statistical consultations go badly when the question is "is this study any good" and well when the question is narrow and tied to a specific appraisal domain.
Do not overlook content expertise. Deciding whether blinding was feasible, whether an outcome measure means what the authors claim, whether a comparator reflects real practice, or whether a co-intervention could plausibly explain a finding requires someone who knows the field. Recruit a clinician, practitioner, or subject specialist onto the team, or arrange for one to review your appraisal decisions after you have drafted them. Where the subject is medical care, this is also the person who can tell you whether a doctor reading your review would find a given deviation from protocol trivial or decisive.
Sometimes the missing information sits with the original study authors. Writing to a corresponding author to ask how the allocation sequence was generated, whether a trial registry entry matches the published outcomes, or how many participants were lost to follow-up is legitimate and common. Give the authors a reasonable deadline, record what you asked and what came back, and rate the domain on the reported evidence if nobody replies. Trial registries, protocol publications, statistical analysis plans, and regulatory documents are other places to look before you settle for an "unclear" or "no information" rating.
Plan for disagreement rather than reacting to it. Decide before appraisal begins who arbitrates when two reviewers cannot converge, whether that is a third reviewer, the review lead, or a methodologist outside the team. Pilot the instrument on a handful of studies, compare judgments, and write down operational definitions for the calls that split your team, so the same reasoning is applied to study fifty as to study five. Keep the supporting quote or page reference alongside every judgment; most disputes turn out to be about what a paper actually said rather than about the instrument. Review software can carry a lot of this load by holding both reviewers' ratings separately, surfacing where they diverge, storing supporting text with each domain, and preserving a record of how conflicts were resolved, which is exactly what peer reviewers and readers later ask about. What no software can do is make the judgment for you, so budget human time for the calls that genuinely require it.
Finally, get outside eyes on your appraisal before the manuscript goes out. Registering a protocol on PROSPERO or a comparable registry forces you to commit to an instrument and a conflict-resolution process in advance. Presenting your risk of bias table at a departmental journal club or methods seminar tends to surface the weak judgments quickly. If a journal's peer reviewers later challenge a rating, the questions are usually answerable if you kept your reasoning and sources attached to each domain from the start, and painful if you did not.
Why Appraising Published Research Matters
Critical appraisal is the step where you stop asking what a study found and start asking whether you should believe it. It is a structured examination of how a piece of research was designed, conducted, analyzed, and reported, aimed at working out how far the reported result could have been shifted by something other than the effect the authors set out to measure. In a systematic review, appraisal sits between screening and synthesis: you have decided a study is eligible, and now you need to decide how much weight its findings can carry.
The reason this matters is simple. A systematic review inherits the flaws of the studies inside it. Pooling results does not cancel out bias; if several trials share the same design weakness, meta-analysis will produce a tighter confidence interval around a number that may be systematically off. A precise summary estimate built from studies with unclear allocation concealment or heavy attrition is more likely to mislead readers than an honest statement that the evidence is thin. Appraisal is what lets you tell those two situations apart, and it is what stops your review from laundering weak primary research into an authoritative-looking conclusion.
It is tempting to treat journal publication as a quality filter and skip ahead. Peer review is genuinely useful, but it varies enormously between journals, it happens before anyone knows how a study will be used, and reviewers see the manuscript rather than the protocol, the raw data, or the analysis code. Published papers can be internally inconsistent, can report outcomes that were not the ones registered, and can describe methods in language vague enough to hide important choices. Appraisal is your own read of the study against the question you are asking, which is almost never the exact question the original authors or their reviewers had in mind.
That last point is worth dwelling on, because it explains why appraisal cannot be outsourced or looked up. Risk of bias is not a property of a paper; it is a property of a specific result used for a specific purpose. A trial might be at low risk of bias for a hard, objectively measured outcome and at high risk for a subjective one assessed by unblinded staff, all within the same publication. A cohort study might be well suited to describing how often something happens and poorly suited to attributing that occurrence to an exposure. If you are appraising, you assess each outcome you plan to use, in the direction you plan to use it, and you record that distinction rather than assigning the study a single overall grade.
What do you actually do with the results? There are four common uses, and it helps to decide which ones you intend before you start. First, appraisal can inform eligibility, if your protocol sets a methodological floor, such as excluding non-randomized designs for a question about effects of an intervention. Second, it can structure the presentation of results, so readers see which findings rest on stronger footing. Third, it drives sensitivity and subgroup analyses: you run the synthesis with and without high-risk studies and report whether the picture changes. Fourth, it feeds directly into certainty-of-evidence frameworks such as GRADE, where study limitations are one of the domains that can lower your confidence in a body of evidence. Appraisal that never gets used for any of these purposes is wasted effort, and reviewers and editors will notice.
Choosing an instrument follows from the design you are appraising and the claim you are making. Randomized trials, non-randomized studies of interventions, diagnostic accuracy studies, prognostic studies, qualitative research, and systematic reviews themselves each have established tools built around the specific ways those designs go wrong. Using a trial checklist on a cohort study, or a generic quality scale on anything, produces scores that look objective and mean very little. Be especially wary of summing domain judgments into a single numeric quality score and then using that score to weight a meta-analysis; the practice is widely criticized because the weights are arbitrary and the domains are not interchangeable.
Appraisal is also where reproducibility is won or lost. The convention in systematic reviews is that two people appraise each study independently, compare judgments, and resolve differences by discussion or by a third reviewer. This is not ceremony. Domain judgments involve interpretation, and two competent reviewers reading the same terse methods paragraph will sometimes disagree about whether outcome assessors could plausibly have known group allocation. What makes the disagreement productive is recording the reason for each judgment: the sentence you relied on, the page or table, and the inference you drew. Support-for-judgment notes let a reader who disagrees with you see exactly where your reasoning diverged from theirs, and they let your own team stay consistent across dozens of studies appraised over several months.
Practically, this argues for capturing appraisal as structured data rather than prose. Domain-level judgments, free-text justifications, the reviewer who made each call, and the date are all things you will need again, whether for a traffic-light figure, a sensitivity analysis, a response to a peer reviewer, or an update of the review in two years' time. Appraisal recorded only in a narrative paragraph has to be redone from scratch every time someone questions it.
What Does a Critical Review Look Like? A Worked Example
A critical review is not a summary with a verdict attached at the end. It is a structured record of where a study's design and conduct could have pushed its results away from the truth, written specifically enough that a second reader can check your reasoning against the paper. The most useful way to see what that means in practice is to walk through one. The example below is illustrative, a hypothetical study, appraised step by step, so that the focus stays on the reasoning rather than on any particular set of results. In your own work, every judgment you see described here would be anchored to a quoted sentence, table, or page number from the paper in front of you.
Start by fixing the question the appraisal serves. Suppose your review asks whether a structured pre-discharge education program, compared with usual discharge instructions, affects readmission within 30 days among adults hospitalized for heart failure. The study in front of you, call it Trial A, is a two-arm randomized trial conducted at a single academic hospital, with readmission ascertained from the hospital's own electronic records. Naming the question matters because appraisal is relative to it. Trial A might be a perfectly sound trial of something else; what you are judging is how much weight its result can carry for your comparison and your outcome.
Next, pick the tool before you read the methods closely, not after. For a randomized trial like this, a domain-based risk-of-bias tool designed for trials is the appropriate instrument; for a cohort study you would reach for a tool built for non-randomized designs, and for a diagnostic accuracy study, something different again. Locking the tool in advance keeps you from unconsciously selecting the instrument that produces the rating you already have in mind. Record the tool and version in your protocol so readers know which lens you used.
Now the domains, one at a time. On randomization, Trial A reports a computer-generated sequence held by the pharmacy and revealed only after consent, that is concealed allocation, and you would note the exact wording that supports the judgment. Suppose the baseline table shows the intervention arm was somewhat older with a higher proportion of prior admissions. A single imbalance in a small trial is not itself evidence of a broken randomization process, and your note should say so explicitly rather than leaving a reader to guess whether you missed it: the sequence generation and concealment are adequately described, one imbalance is present, and it is plausibly chance. That is a low-concern judgment with a documented caveat.
On blinding and deviations from intended interventions, an education program cannot be masked from patients or the nurses delivering it. The question is not whether blinding was achieved but whether the lack of it plausibly affected this outcome. Readmission decisions are made partly by clinicians who knew which arm the patient was in, and patients who received extra attention may present differently to the emergency department. That is a real mechanism, and it belongs in the appraisal note. If the trial reports an intention-to-treat analysis including everyone as randomized, you record that as a mitigating feature.
On missing outcome data, ascertainment from the hospital's own records is a specific vulnerability: patients readmitted to a different hospital are invisible. Suppose Trial A does not report whether it checked a regional or state-wide database. That is not a small technicality, it is a mechanism by which outcomes go missing, and it may operate unequally if the intervention changed where patients chose to seek care. Your note should say that missingness is likely non-trivial, that its direction is uncertain, and that the paper gives you no way to bound it.
On measurement of the outcome, administrative readmission is relatively robust to observer judgment once it is captured, which distinguishes it from an outcome like "adherence" scored by an unblinded coordinator. On selection of the reported result, you check the registry entry. If the registration lists a composite of readmission or death as the primary endpoint but the published paper leads with readmission alone, you flag a possible selective-reporting concern and note the discrepancy precisely, including the registration number and the date the record was last updated. If no prospective registration exists, that itself is what you record.
The write-up is the part most reviewers shortchange. A usable critical review of Trial A reads something like: Randomized, single-center trial with adequate sequence generation and concealed allocation. Participants and care providers were unblinded, which is unavoidable for this intervention but plausibly relevant because readmission decisions involve clinician discretion. Outcomes were ascertained from the study hospital's records only, with no described linkage to external data; the extent and direction of resulting misclassification cannot be determined from the report. The published primary outcome differs from the registered composite endpoint. Overall: high concern for bias, driven by outcome ascertainment and by the discrepancy with registration. Notice that the paragraph gives the driver of the overall rating, not just the rating. Notice too that it commits to "cannot be determined" where the paper is silent, instead of guessing.
Two people should do this independently, and the disagreements are the valuable part. In practice, most disagreements come from one reviewer having found a sentence the other missed, an appendix line about database linkage, a protocol amendment, rather than from genuine differences of judgment. Reconciliation therefore works best as a conversation about evidence in the paper, with a third reviewer available for the residual cases where two people read the same sentence differently. Keep a short log of how each unresolved disagreement was settled; it takes minutes and it answers the peer reviewer who asks how you handled them.
Appraisal software earns its place at exactly these friction points. It should hold the tool's signaling questions next to a free-text field for supporting quotes, keep the two reviewers' answers separated until both are submitted, surface the specific domains where they diverge, and carry the rationale text through to the figure so the traffic-light plot in your manuscript is backed by the sentences that produced it. When appraisal lives in a spreadsheet, the ratings usually survive and the reasons usually do not, and the reasons are what make the review critical rather than merely rated.
How to Write an Appraisal Example of a Study
A written critical appraisal is not a grade or a verdict; it is a short, auditable argument about how much confidence a reader should place in a study's results, and why. The mistake most new reviewers make is writing an opinion ("this study was poorly conducted") without the reasoning that produced it. A usable appraisal example does the opposite: it names the domain being judged, quotes or paraphrases the specific detail from the paper that drove the judgment, states the judgment, and explains which direction the problem would push the result. Anyone reading it later, a coauthor, a peer reviewer, an editor, or you in eight months, should be able to reconstruct your thinking without reopening the full text.
Start by fixing the unit you are appraising. This sounds pedantic, but it prevents most of the confusion in appraisal write-ups. You are almost never appraising "the study" as a whole; you are appraising a specific result for a specific outcome as it applies to your review question. The same trial can be rated one way for a primary outcome measured by blinded assessors and another way for a self-reported secondary outcome collected only at long-term follow-up. Write the outcome and time point into the first line of your appraisal so the scope of your judgment is unambiguous.
Next, work domain by domain rather than writing free prose. Whichever instrument you are using, a risk-of-bias tool for randomized trials, a tool for non-randomized studies of interventions, an appraisal checklist for diagnostic accuracy, prognosis, or qualitative work, it gives you the headings. Use them. Under each heading, write two things: the evidence from the paper (the "support for judgment") and the resulting rating. Keep them physically separate in your notes or your software fields. Reviewers who blend the two end up with paragraphs that assert a rating and then restate it in different words, which is unfalsifiable and useless at consensus stage.
For the support-for-judgment text, prefer the paper's own words where a phrase is doing real work. If the methods say participants were assigned "using a computer-generated random sequence held by an off-site pharmacist," quote it, that single clause resolves both sequence generation and allocation concealment. If the paper says only "patients were randomized," say that, and say plainly that no further detail was reported. The distinction between a method that was inadequate and a method that was not described matters enormously, and it is the distinction most often lost in sloppy write-ups. When you have gone looking for the missing detail, checking a protocol, a trial registry entry, a statistical analysis plan, or a prior publication of the same cohort, record that you looked and what you found or did not find.
Then write the judgment sentence. Use the categories your instrument defines, spelled exactly as the instrument spells them, so that your text and your data table cannot drift apart. Follow the rating with a direction-of-effect sentence wherever you can reason about it: attrition concentrated in one arm, unblinded assessment of a subjective outcome, or a change in the primary outcome between registration and publication all have plausible directions, and naming the direction is what turns an appraisal into something a synthesis can actually use. If you genuinely cannot tell which way a limitation cuts, write that instead of implying a direction you have not reasoned through.
Here is the shape of a worked appraisal for an illustrative, hypothetical trial. Unit: mortality at 30 days. Randomization: "block randomization with sealed opaque envelopes prepared centrally", adequate sequence generation and concealment; low concern. Blinding: participants and clinicians aware of assignment, but the outcome is all-cause death ascertained from records, so lack of blinding is unlikely to matter for this outcome; low concern. Missing data: 6 of 240 participants lost, distributed 4 and 2 across arms, with reasons given and an intention-to-treat analysis reported; low concern. Selective reporting: registry entry lists 30-day mortality as primary, matching the publication; low concern. Overall for this outcome: low concern, with the note that the same trial's quality-of-life outcome would be appraised separately and more cautiously given the open-label design and self-reported instrument. Notice that every clause either reports something from the paper or draws an inference explicitly from it.
Length should follow function. In a full systematic review, most domains need one to three sentences; a domain where you and your co-reviewer disagreed, or where the judgment is doing heavy lifting in the synthesis, deserves a longer note explaining the resolution. For teaching examples or journal club write-ups, longer narrative appraisals are appropriate because the reasoning itself is the product. What does not vary is the structure: evidence, judgment, consequence. If a paragraph contains no detail traceable to the study, it is commentary, not appraisal.
Write independently before you compare. Two reviewers appraising the same study should produce their support-for-judgment text without seeing each other's, then reconcile. Keep both original entries and a short consensus note recording what changed and on what basis, this is the part that peer reviewers and updaters most often ask about, and it is trivially easy to capture at the time and painfully hard to reconstruct later. If you are appraising at any scale, pilot the process on three to five studies first, compare your wording, and tighten the decision rules you are applying before the disagreements multiply across the whole set. Store the appraisal text in a structured field tied to the study record and the outcome, not in a comment thread or a loose document, so it can be exported straight into a risk-of-bias figure, an appendix table, or the narrative that explains how appraisal shaped your conclusions.
Finally, avoid the recurring failure modes: converting a checklist into a numeric score and reporting the total instead of the domain judgments; appraising reporting quality when you meant to appraise study conduct; letting the plausibility of a study's findings influence the rating of its methods; and writing "unclear" as a resting place rather than as a documented conclusion after looking for the information. An appraisal example is well written when a stranger can read it, disagree with your rating, and point precisely to the sentence where they would have decided differently.
Examples of Critical Writing About Research Evidence
Critical writing is easiest to learn by contrast. Descriptive writing tells the reader what a study did and what it reported; critical writing tells the reader what that report is worth, and why. The examples below use placeholder interventions and outcomes ("Intervention A," "the primary outcome") rather than real figures, because the point being illustrated is the sentence structure and the reasoning move, not the numbers. When you write your own appraisal, the numbers come from the papers in front of you.
Appraising a single study. A purely descriptive sentence reads: "Smith et al. conducted a randomized trial of Intervention A versus usual care in adults recruited from two outpatient clinics, and reported a difference favoring Intervention A on the primary outcome." The critical version keeps that information but attaches a judgment to it: "Smith et al. randomized adults from two outpatient clinics to Intervention A or usual care and reported a difference favoring Intervention A. Because outcome assessors also delivered the intervention and the primary outcome was clinician-rated, the estimate is vulnerable to detection bias in the direction of the reported effect. The trial also reports a per-protocol analysis only, with attrition described in the text but not accounted for in the estimate, so the reported difference is best read as the result among participants who completed the intervention rather than among all those who began it."
Notice what the second version adds: a named bias domain, the mechanism by which it operates, the likely direction of its influence, and a restatement of what the finding actually supports. Those four moves, domain, mechanism, direction, restatement, are the backbone of most good appraisal sentences.
Appraising applicability rather than validity. A study can be internally sound and still be a poor fit for your question. Critical writing separates the two: "The trial's randomization, allocation concealment, and outcome ascertainment are well reported, and we judged risk of bias low across domains. Applicability to our review question is more limited: participants were recruited from a specialty referral center, comorbid conditions were exclusion criteria, and the comparator was a usual care pathway with more frequent follow-up contacts than is typical of the primary care setting our question addresses. We therefore treat this study as directly informative about efficacy under favorable conditions and only indirectly informative about the effectiveness question our review poses."
Appraising a body of evidence. Appraisal at the review level asks different questions: does the collection of studies hang together, and does it cover the question? For example: "The four included trials point in the same direction, but they differ in ways that make pooling difficult to interpret. Two used a self-reported outcome measure and two used a performance-based measure; follow-up ranged from immediately post-intervention to several months later; and the intervention was delivered by trained specialists in three trials and by existing staff in the fourth. The statistical heterogeneity we observed is therefore unsurprising, and we present the results narratively rather than as a single pooled estimate. Subgroup exploration by delivery model was planned but is underpowered with four studies and should be read as hypothesis-generating."
Writing about reporting and publication bias. Avoid both silence and overstatement here: "Three of the six included trials were registered prospectively; of these, two changed the primary outcome between registration and publication without explanation. We could not locate results for two registered trials that met our eligibility criteria and whose completion dates have passed, and our attempts to contact the investigators went unanswered. With six studies, funnel plot asymmetry cannot be assessed meaningfully. We note selective reporting as a plausible concern that we were unable to quantify, rather than as a bias we have measured."
Writing about your own review's limitations. Critical writing includes turning the appraisal lens inward, specifically and without ritual apology: "Our search was limited to English-language records, which may have excluded relevant trials from regions where this intervention is more commonly delivered. Screening was performed in duplicate, but data extraction for two studies was completed by a single reviewer with verification only of the primary outcome. We prespecified a random-effects model and did not deviate from it; however, our decision to combine two related outcome measures into a single construct was made after seeing the extracted data and is documented as a post hoc change in the protocol amendment log."
Language that carries the judgment. Critical writing depends heavily on verb and hedge choice. "The study demonstrates" claims more than "the study reports"; "is consistent with" claims less than "shows"; "we judged" signals that a person made a decision that another appraiser might make differently. Say "the estimate may be biased toward the intervention" rather than "the study was biased," because bias is a property of an estimate under particular conditions, not a verdict on the authors. Attribute judgments to yourself and your co-reviewers, since a reader who disagrees needs to know whose reasoning to argue with.
Common failure modes. Three patterns weaken otherwise careful appraisals. The first is the tool dump: reproducing every domain rating from your appraisal instrument in prose without saying which limitations actually matter for your question. The second is the vague dismissal, "the included studies were of generally poor quality", which gives the reader no way to evaluate or reuse your reasoning. The third is the disconnected limitations paragraph, where problems are acknowledged at the end of a section but the conclusion is written as if none of them existed. The test for the last one is simple: read your conclusion sentence, then read your appraisal, and ask whether the confidence in the first survives the second. If it does not, revise the conclusion rather than softening the appraisal.
Sentence Starters for Writing Critical Analysis
A sentence starter is scaffolding, not a substitute for judgment. The reason appraisal write-ups read thinly is rarely that the author lacked vocabulary; it is that the sentence stopped at the verdict. The pattern worth practicing has three parts: name what you observed in the study, name the mechanism by which it could distort the result, and state what that means for how much weight the finding carries. "The trial was low quality" does none of those things. "Outcome assessors were aware of allocation, and because the primary outcome was a clinician-rated symptom scale, ratings could have favored the intervention arm" does all three. The starters below are grouped by the job they do, and each one is written to invite a reason after it.
To describe before you judge. Keep the neutral, descriptive sentences separate from the evaluative ones, so a reader can check your reasoning against the study rather than taking your conclusion on faith. Useful openings: "The authors randomized participants to…", "Allocation was concealed by…", "The primary outcome was measured as…", "Follow-up was reported for X of Y randomized participants…", "The analysis population was defined as…", "The comparator in this study was…", "The study reports outcomes at a single time point, at…". These are dull on purpose. Description that a skeptical reader can verify is what earns your later judgments any credit.
To name a bias mechanism and its likely direction. This is where most write-ups need the most help. Try: "Because participants selected their own group, differences at baseline in… could account for part of the observed difference"; "Attrition was uneven across arms, with more loss in the… group, which could bias the estimate toward…"; "The outcome was self-reported and participants were aware of their assignment, so expectation effects cannot be separated from the treatment effect"; "The protocol specified… but the paper reports…, which raises the possibility of selective outcome reporting"; "Adjustment was made for…, but no adjustment is reported for…, which is plausibly associated with both exposure and outcome"; "The direction of this bias, if present, would most likely inflate the observed benefit"; "It is not clear from the report whether…, and the authors do not describe…". That last one matters: unclear reporting and demonstrated flaws are different findings, and saying so is more accurate than guessing.
To address applicability and indirectness. A methodologically sound study can still be a poor fit for your question. Openings that force the comparison: "The population enrolled differs from the review's target population in that…"; "The intervention was delivered by… under conditions that differ from routine practice in…"; "The comparator was…, rather than the alternative most relevant to this question, which is…"; "The outcome measured was…, which is a surrogate for the outcome of interest, namely…"; "Follow-up extended to…, which is shorter than the horizon over which this decision plays out"; "Applying this estimate to… requires the additional assumption that…".
To talk about precision, consistency, and the shape of the body of evidence. Try: "The confidence interval spans values that would support opposite decisions, from… to…"; "The estimate is compatible with both a meaningful benefit and no meaningful difference"; "The point estimates across studies fall in the same direction, but the magnitude varies from… to…"; "The two studies disagree, and the difference in… offers one candidate explanation"; "The pooled estimate is dominated by a single large study, so the summary reflects that study's design choices more than the others'"; "Because only… studies contribute to this comparison, the consistency of the finding cannot be assessed"; "This finding rests on a single study and should be read as preliminary until replicated".
To state your own certainty honestly. Hedging is only useful when the hedge says why. Prefer "Our confidence in this estimate is limited mainly by…" over "more research is needed." Other options: "The available evidence supports…, with the caveat that…"; "We interpret this result cautiously because…"; "Two readings of these data are defensible: … or …"; "The evidence is insufficient to distinguish between… and…"; "Nothing in these studies speaks to…". And for your own review: "Our search was limited to…, so studies published in… may be missing"; "Screening decisions for… were judgment calls, and reasonable reviewers could have decided differently"; "We chose to pool… despite…, and results from an analysis that separates them are reported below".
Two habits that matter more than the phrasing. First, anchor evaluative sentences to whatever tool you are using, so your prose and your domain judgments tell the same story; if you write that blinding was adequate, the corresponding domain rating should not say otherwise. Second, watch for starters that have hardened into filler. "It should be noted that", "the study suffers from several limitations", and "the results should be interpreted with caution" can all be deleted without losing information. Replace each with the specific limitation, the specific mechanism, and the specific consequence for the reader's decision. If you cannot fill in those specifics, that is a signal to reread the paper rather than reach for another phrase.
How to Perform a Critical Evaluation of a Study
Critical appraisal is the step where you decide how much weight a study's results deserve. It is not a verdict on whether the authors are competent, and it is not a summary of whether you like the findings. The question you are answering is narrower and more useful: given how this study was designed, conducted, and reported, how likely is it that its estimate of the effect differs from the truth, and in which direction? Everything below is a way of making that judgment explicit enough that another reviewer could follow your reasoning and disagree with it on specific grounds.
Start by pinning down what the study actually is before you appraise it. Identify the design as conducted rather than as labeled in the title, since papers described as trials are sometimes single-arm cohorts and papers described as cohorts are sometimes cross-sectional. Note the unit of randomization or sampling, the comparator, the primary outcome as prespecified, and the time points. Then choose an appraisal instrument that matches that design. Randomized trials, non-randomized studies of interventions, diagnostic accuracy studies, prognostic studies, qualitative research, and systematic reviews all have their own established tools, and applying a trial-oriented checklist to an observational study produces judgments that look rigorous but mean very little. Pick the tool during protocol development, not after you have read the papers, so that your choice is not shaped by the results you happen to have seen.
Read the paper in an order suited to appraisal rather than the order it is printed in. Methods first, then the flow of participants, then the tables, then the results text, and the discussion last. Reading the abstract and discussion early tends to anchor you to the authors' interpretation, which is exactly what you are trying to evaluate independently. Keep the registration record, protocol, and any supplementary appendices open alongside the article; much of what determines a risk-of-bias judgment, the randomization procedure, the prespecified outcome list, the analysis plan, lives outside the main text.
Work through the domains one at a time and make a judgment on each before forming any overall impression. The domains vary by tool, but the underlying questions are consistent. How were groups formed, and could the way participants entered the study have made them systematically different at baseline? For non-randomized studies, which confounders matter for this question, were they measured, and were they measured before the exposure? Who knew which group a participant was in, and could that knowledge have influenced how care was delivered or how outcomes were recorded? How was the outcome measured, was it the same instrument and same schedule in both groups, and is subjective assessment involved? How many participants are unaccounted for at the end, do the numbers in the flow diagram reconcile with the numbers in the analysis tables, and is there reason to think dropout is related to the outcome? Finally, does the set of results reported match the set that was planned, or have outcomes, time points, or subgroups appeared and disappeared between protocol and publication?
Record each judgment with support, not just a label. For every domain, write down the answer you reached, a short quotation or a specific location in the paper, and one sentence on why that led to your rating. If the paper is silent on a domain, mark it as no information rather than as low risk; absence of a description is not evidence of good practice, and it is also not evidence of bad practice. Where a problem exists, note the likely direction of the resulting bias if you can reason about it, since "this would tend to exaggerate the difference between groups" is far more usable later than "some concerns." This documented trail is what allows the appraisal to feed a sensitivity analysis, a certainty-of-evidence assessment, or a straightforward paragraph in your results narrative.
Have two people appraise each study independently, then compare. Disagreements are informative: most of them come from one reviewer having found a detail in a supplement, from different readings of an ambiguous sentence, or from unstated differences in how strictly a domain is being applied. Resolve them by discussion, bring in a third reviewer when discussion stalls, and when you find yourselves applying a rule that is not written down, write it down and apply it to the studies you have already done. A short decision log of these calibration rules is worth more to consistency than any amount of reviewer experience held in someone's head.
Two habits are worth avoiding. The first is collapsing domain judgments into a numeric score and then ranking or excluding studies by that score; a total treats unrelated flaws as interchangeable and hides which specific problem is driving your concern. Report the domains. The second is conflating internal validity with relevance. A well-conducted study in a population, dose, or setting unlike yours is not biased, it is simply less applicable, and the two issues should be recorded separately so that neither is used to dismiss the other. Once appraisal is complete, use it: decide in advance whether high-concern studies will be excluded from the primary synthesis, restricted to a sensitivity analysis, or retained with their limitations discussed, and state that decision in the protocol so the choice is not made after seeing which way the results fall.
Key Elements Every Critical Review Should Include
A critical review is not a summary with an opinion attached. It is a structured judgment about whether a study's design, conduct, and reporting support the conclusions its authors drew, and whether those conclusions are useful for your question. Whether you are appraising a single paper for a journal club or a hundred records inside a systematic review, the same core elements need to be present and visible to anyone reading your work later.
A clear statement of the question you are appraising against. Appraisal is always relative to a purpose. The same randomized trial can be well conducted and still be nearly useless to you if it enrolled a population, used a comparator, or measured an outcome that does not match your review question. Write down the question, population, intervention or exposure, comparator, outcome, and setting, before you start reading, and appraise against it. Reviewers who skip this step tend to drift into appraising the study on its own terms rather than yours.
An accurate description of what the study actually did. Before judging anything, record the design as conducted rather than as labeled. Papers describe themselves as randomized, prospective, or matched in ways that the methods section does not always support. Capture the unit of allocation and analysis, how participants entered the study, what the comparison group was, how long follow-up ran, how the outcome was defined and measured, and how many participants were unaccounted for at the end. This descriptive layer is what your judgments will hang on, and it is what makes your appraisal auditable.
Assessment against a named, pre-specified tool. Choose the appraisal instrument before screening ends, match it to the design, and name it in your methods. Different tools ask different questions, domain-based risk-of-bias instruments for trials, tools built around selection and confounding for cohort and case-control studies, tools with items on sampling and reflexivity for qualitative research, and separate instruments for diagnostic accuracy, prevalence, economic evaluations, and existing reviews. Mixing tools within a single body of evidence is sometimes unavoidable when designs differ, but decide in advance which tool applies to which design and say so.
Judgments recorded at the domain level, not just an overall verdict. The useful output of appraisal is a set of specific findings: how the allocation sequence was generated, whether outcome assessors could have known group assignment, whether the analysis population matched the randomized population, whether confounders were identified and handled, whether the outcome measure was validated. A single global rating collapses all of that into one label and hides the reasoning. Keep the domains separate, then derive an overall judgment from them if your synthesis method calls for one, and state the rule you used to derive it.
Supporting evidence for each judgment. Every domain rating should carry the quote, table reference, or page number that justifies it, plus a sentence of reasoning where the judgment required interpretation. This is the element most often skipped and the one that causes the most rework. Without it, disagreements between reviewers cannot be resolved except by rereading the paper, peer reviewers cannot check your work, and updates to the review months later start from scratch. Most review management software provides a free-text support field beside each domain response; treat it as required rather than optional.
A distinction between what was not done and what was not reported. These are different problems and they deserve different labels. A study that used a valid method but described it poorly is a reporting failure; a study that omitted the method is a conduct failure. Many tools handle this with an explicit "no information" or "unclear" category. Use it, and note whether you attempted to resolve the gap, through a protocol, a trial registry entry, a supplement, or contact with the authors.
The direction and plausible consequence of each concern. Saying a study has high risk of bias in a domain is less informative than saying which way that bias would be expected to push the result and whether it could plausibly account for the finding. This turns appraisal from a scoring exercise into an interpretive one and is what lets you use appraisal in synthesis, for subgroup or sensitivity analyses restricted to studies without a given concern, or for narrative explanation of heterogeneity.
An applicability judgment kept separate from internal validity. Indirectness, generalizability, and fit to your question are real limitations, but they are not the same as bias, and merging them produces ratings nobody can interpret. Record them in their own field.
The process itself. State how many people appraised each record, whether they worked independently, how disagreements were resolved, whether appraisers were trained or calibrated on a pilot set, and whether anyone appraised a study they had authored. If you piloted the tool and revised your decision rules, say what changed.
A presentation of results that a reader can inspect. Report appraisal findings per study and per domain, not only in aggregate, and make the supporting text available in a supplement or appendix. Aggregate figures such as domain-level summary plots are helpful as an overview but should not be the only thing you publish. Finally, note the limitations of the appraisal itself, tools not designed for the designs you encountered, single-reviewer appraisal where resources were short, or judgments made from abstracts alone. Being explicit about where your appraisal is weak is part of doing it well.
Another Word for Critical Evaluation: Common Synonyms
The most common synonym for critical evaluation in evidence synthesis is critical appraisal. If you are writing a protocol, a methods section, or a training document for a review team, that is the term most readers in the field will expect. Other words and phrases that get used for the same broad activity include quality assessment, methodological assessment, evidence appraisal, validity assessment, risk of bias assessment, study appraisal, and, in more general or academic writing, critique, scrutiny, examination, or simply assessment.
Those terms are not all equivalent, and treating them as interchangeable is where a lot of methods sections get muddy. "Critical appraisal" is the umbrella: the structured judgment of whether a study's design, conduct, and reporting support the conclusions drawn from it. "Risk of bias assessment" is narrower, it asks specifically whether features of the study's design or conduct could have pushed the result away from the truth, and it is usually answered domain by domain with a named instrument. "Quality assessment" is looser still, and different authors use it to mean risk of bias, reporting completeness, or an unspecified blend of both. If a reviewer or editor asks you to clarify your terminology, this ambiguity is usually why.
Current methodological guidance generally pushes authors away from the vague "study quality" and toward the more specific "risk of bias," on the grounds that a study can be well reported and still be at high risk of bias, or poorly reported and yet soundly conducted. Cochrane's guidance and the PRISMA reporting items both frame the relevant step in terms of risk of bias in the included studies rather than their overall quality. Adopting that vocabulary makes your methods easier to check, because "risk of bias" points to a specific instrument and a specific set of domains, whereas "quality" does not.
For practical purposes, pick your term based on what you actually did and then use it consistently. If you applied a domain-based instrument to each included study, say you assessed risk of bias and name the tool. If you used a checklist that mixes reporting, design, and analysis items, critical appraisal is the safer and more honest label. If you rated the body of evidence for an outcome rather than individual studies, that is a different exercise again, the usual wording is certainty of the evidence or strength of the evidence, and calling it critical appraisal will confuse readers who know the distinction.
Audience matters too. In a journal manuscript, a grant methods section, or a review protocol, use the technical term. In a plain-language summary, a stakeholder briefing, or a public-facing report, phrases like "we checked how reliable each study was" or "we weighed how much confidence the studies deserve" communicate better than "risk of bias assessment," which reads as jargon to non-specialists. Some teams write both: the technical phrase in the methods, the plain phrase in the abstract or summary.
One habit that saves time later: define your term once, early, and attach it to the instrument you used. A single sentence, naming the activity, the tool, who did it, and how disagreements were resolved, removes most of the ambiguity that synonym choice introduces. Whatever word you land on, the reader's real question is not what you called the step but whether they can reconstruct what you judged and how.