How to Conduct a Systematic Review: A Reproducible Step-by-Step Workflow

Written by Elena BrooksLast updated: August 12, 202610 min read

Featured Snippet Answer
A systematic review answers a focused question by using predefined, transparent methods to search for, select, appraise, and synthesize all eligible evidence. Unlike a narrative overview, its protocol, eligibility rules, search strategy, screening decisions, extraction process, and synthesis methods must be documented well enough for others to evaluate or reproduce.
Decision entry: Choose a systematic review only when the question, evidence base, team, time, and reporting requirements justify a comprehensive and auditable process.
Workflow: question → protocol → eligibility → search → deduplicate → screen → extract → appraise → synthesize → report.
Research smarter with Acade
One Academic Agent for literature search, research design, writing, and more.
Try Start free →Systematic, Narrative, or Scoping Review?
Review type | Best for | Defining feature |
|---|---|---|
Systematic review | focused answer or effect question | predefined, reproducible methods and critical appraisal |
Scoping review | mapping concepts, evidence types, or gaps | broad charting of a field |
Narrative review | interpretive overview | flexible selection and synthesis |
A systematic review is not automatically better. A question that is too broad, a deadline of several days, or a single untrained reviewer may make a formal systematic review infeasible.
Systematic Review vs Meta-Analysis
A meta-analysis statistically combines compatible quantitative results. A systematic review may include a meta-analysis, but it may instead use structured narrative synthesis when studies differ too much in participants, interventions, outcomes, designs, or measures. Conversely, pooling conveniently available studies without a systematic search is not a systematic review.
Decide Whether a Systematic Review Is the Right Design
Use a systematic review when a focused question requires a comprehensive, transparent assessment of existing evidence. Do not select the label merely because it sounds rigorous.
Decision question | Systematic-review signal | Consider another approach when |
|---|---|---|
Is the question focused? | Population, intervention/concept, comparator, and outcome can be specified | The aim is to map terminology or study types |
Is exhaustive searching necessary? | Missing eligible studies could change the answer | A selective conceptual argument is intended |
Are methods auditable? | Criteria and decisions can be documented in advance | Selection will remain informal |
Are resources available? | Databases, full texts, trained reviewers, and time are available | One person has only a few days |
Is synthesis meaningful? | Studies can be grouped using prespecified logic | The field is too broad for coherent grouping |
The output is a justified design decision recording purpose, scope, likely evidence, team, timeline, and reporting standard.
Choose a Framework Without Forcing It
PICO is often useful for intervention questions: Population, Intervention, Comparator, and Outcome. PICo may support qualitative questions by specifying Population, phenomenon of Interest, and Context. SPIDER can help frame Sample, Phenomenon of Interest, Design, Evaluation, and Research type. These are aids, not universal rules.
Input: the decision the review should inform and disciplinary conventions. Output: a structured question plus operational definitions. Common error: placing vague terms such as “adults,” “technology,” or “well-being” into a framework without defining them. Done when: reviewers can apply the concepts consistently to candidate records.
Build the Protocol Before Seeing Results
A protocol should state the rationale, question, eligibility rules, information sources, search approach, screening method, extraction fields, appraisal method, synthesis plan, and amendment procedure. Cochrane explains that specifying methods in advance reduces the chance that decisions are shaped by study findings (Handbook Chapter 1). Check whether a registry, funder, institution, or journal expects prospective registration.
Decision | Example operational rule | Audit evidence |
|---|---|---|
Population | adults with clinician-diagnosed insomnia | eligibility manual |
Intervention | app-delivered CBT with defined core components | coding guide |
Comparator | usual care, waitlist, or active control | comparator categories |
Outcome | prespecified validated sleep measures | outcome dictionary |
Design | randomized or controlled clinical studies | design rule |
Amendments may be necessary. Date them, explain them, and state whether emerging results influenced the change.
Step-by-Step Review Workflow
1. Frame the question and protocol
Use a framework appropriate to the question: PICO for many intervention questions, PICo for some qualitative questions, or SPIDER for certain qualitative and mixed-method evidence. Define outcomes, study designs, settings, dates, languages, and publication types before screening. Record any justified protocol changes.
2. Design a reproducible search
Translate each concept into keywords, synonyms, subject headings, spelling variants, and database syntax. Search sources appropriate to the discipline; one general search engine is rarely enough. Save the exact query, platform, date, filters, and result count. Cochrane emphasizes thorough, objective, reproducible searching and appropriate information-specialist involvement (Handbook Chapter 4).
3. Deduplicate and screen
Preserve a master library, remove duplicate records without losing provenance, and screen titles/abstracts before eligible full texts. Apply the same criteria to every record and record one defensible exclusion reason at full-text stage. Independent duplicate screening is often preferred where resources permit.
4. Extract and appraise
Pilot a structured extraction form. Capture study identifiers, design, sample, setting, intervention/exposure, comparator, outcomes, results, limitations, and reviewer notes. Use a risk-of-bias or quality tool suited to the study design; do not invent a universal score.
5. Synthesize and report
Group studies according to the protocol, describe heterogeneity, and distinguish absence of evidence from evidence of no effect. PRISMA 2020 provides a reporting checklist and flow-diagram templates; it is a reporting guideline, not a substitute for sound conduct (PRISMA 2020).
Search, Screening, and Extraction Artifacts
Create a search-concept table before writing database syntax. For each concept, record controlled vocabulary, free-text variants, spelling forms, and exclusions. Document the database and platform because syntax differs. “We searched PubMed with relevant keywords” is not reproducible.
Pilot eligibility criteria on a sample, discuss disagreements, and revise ambiguous rules before full screening. Preserve a master library, deduplication history, title/abstract decisions, full-text decisions, and one operational exclusion reason per excluded full text. “Not relevant” is not sufficient.
The extraction form should separate reported facts from reviewer calculations and judgments. Cochrane recommends planning data in advance, piloting the form, and using more than one person for critical outcome extraction because errors may be difficult to detect later (Handbook Chapter 5). Select risk-of-bias methods appropriate to each study design rather than inventing a universal quality score.
Decide How to Synthesize
Before meta-analysis, check clinical, methodological, and statistical compatibility. If pooling is not defensible, use structured synthesis that explains grouping, direction, magnitude, certainty, and limitations. Do not count how many studies are statistically significant and call that synthesis.
PRISMA 2020 supplies reporting checklists and flow-diagram templates, but it does not certify review conduct (PRISMA 2020). Readers should be able to reconstruct what was searched, how records moved through screening, what was included, and how conclusions followed from the evidence.
Minimum Audit Trail
Retain the dated protocol and amendments, complete searches, exported records, deduplication log, screening decisions, exclusion reasons, extraction forms, appraisal judgments, analysis files, and PRISMA materials.
Failure state: an unexplained criterion change or untraceable exclusion. Recovery: stop synthesis, reconstruct the record where possible, document uncertainty, and repeat affected steps. Done when: an independent reviewer can follow the decision chain from question to conclusion without undocumented memory.
Complete Worked Example
Illustrative question: For adults with chronic insomnia, do app-based cognitive behavioral interventions improve validated sleep outcomes compared with usual care?
The protocol defines adults, eligible intervention features, comparators, validated outcomes, controlled designs, databases, dates, and synthesis rules. During screening, a mindfulness-only app is excluded because it does not meet the intervention definition. If outcome scales and follow-up periods are incompatible, the team may use structured synthesis rather than forcing a pooled estimate.
Worked protocol excerpt
Population: adults aged 18 or older with chronic insomnia defined by an operational diagnostic rule. Intervention: a mobile application delivering defined CBT-I components. Comparators: usual care, waitlist, attention control, or another eligible intervention. Outcomes: prespecified validated measures at the protocol-defined follow-up. Design: controlled comparative studies. These illustrative rules require clinical and methodological review.
Worked search and screening decisions
Concept | Illustrative variants | Decision still required |
|---|---|---|
insomnia | insomnia, chronic insomnia, sleep initiation disorder | diagnostic boundaries |
mobile application | smartphone app, mobile app, digital therapeutic | whether web-only programs qualify |
CBT-I | cognitive behavioral/behavioural therapy for insomnia | minimum eligible components |
An information specialist translates these concepts into each database’s controlled vocabulary and syntax. During screening, an eligible randomized app trial advances; a general sleep-hygiene messaging study is excluded if it lacks the defined intervention; and multiple reports from one trial are linked rather than counted as separate studies.
Worked extraction and synthesis decision
The team records design, sample, diagnostic criteria, intervention, comparator, outcomes, missing-data handling, results, funding, and conflicts. Risk-of-bias judgments retain supporting locators. If studies differ in intervention components, scales, and follow-up, the team first groups them using prespecified clinical logic and then determines whether compatible effect measures and a defensible model exist. If not, it reports a structured synthesis and explains why pooling was rejected.
This illustrative example is complete when every inclusion, extraction, appraisal, and synthesis decision traces to an advance rule or documented amendment.
Common Systematic-Review Failures
calling a database search and narrative summary “systematic” without a protocol;
searching one convenient source without justification;
changing eligibility after seeing preferred results;
losing links between reports and underlying studies;
treating PRISMA compliance as proof of low bias;
combining quality scores mechanically;
equating a non-significant result with evidence of no effect;
concluding that more research is needed without specifying the remaining uncertainty.
For clinical questions, this workflow does not replace subject expertise, statistical review, information-specialist input, or institutional and journal requirements.
How Acade Helps
Acade can help refine the question, organize search concepts, create screening and extraction structures, summarize user-supplied papers, and draft an evidence-linked outline. It cannot guarantee search completeness, independently adjudicate eligibility, perform valid risk-of-bias judgments without review, or turn a narrative workflow into a compliant systematic review.
Prompt: “Using this protocol draft, build a search-concept table and extraction template. Mark every eligibility, appraisal, and synthesis decision that requires reviewer confirmation.”
Recommended Page Experience
Module | Input | Output | Error state/mobile |
|---|---|---|---|
Review-type selector | purpose, scope, deadline | recommended review type | warns when resources conflict; card layout |
Framework selector | question elements | PICO/PICo/SPIDER worksheet | flags missing outcome/context |
Screening workflow | citations + criteria | decision log | requires reason; swipe disabled to prevent accidents |
Extraction table | included study | structured fields | highlights missing provenance |
These are recommended page experiences unless verified in the live product.
Quality Checklist
Protocol and eligibility criteria precede screening.
Searches, platforms, dates, and restrictions are reproducible.
Duplicate removal and exclusion reasons are documented.
Extraction and appraisal are piloted and checked.
Synthesis matches clinical/methodological compatibility.
PRISMA items are reported where applicable.
UI Proof, Links, and Schema
Show a genuine Acade screenshot of a search-concept or evidence-table workflow. Alt: “Acade workspace organizing systematic-review search concepts and extraction fields for reviewer verification.” Link to /features/ai-literature-search, /features/literature-review, and /features/research-design. Use Article, HowTo, and FAQPage schema. Verify live features, protocol registry requirements, reporting extension, sources, and journal rules before publishing.
Human Verification Items
Confirm the protocol, eligibility criteria, database coverage, exact searches, screening decisions, extraction accuracy, risk-of-bias judgments, synthesis method, PRISMA reporting, and current Acade capabilities.
FAQ
How long does a systematic review take?
It depends on scope, team size, evidence volume, access, and methods; a rigorous review commonly requires substantial time and more than one reviewer.
Does every systematic review need meta-analysis?
No. Pooling is appropriate only when data and studies are sufficiently compatible and the planned model is defensible.
Is PRISMA a method for conducting reviews?
PRISMA primarily guides transparent reporting. Conduct should follow an appropriate methodological handbook and protocol.
Can Acade screen studies automatically?
It may support organization and summarization, but accountable reviewers must confirm eligibility and resolve uncertainty.

About the author
Elena Brooks
Academic Research Content Editor at Acade
Elena Brooks is an Academic Research Content Editor at Acade. She creates practical, evidence-informed content about literature research, research design, academic writing, and the responsible use of AI in scholarly work. She works with Acade’s product team to evaluate research workflows, verify product capabilities, and translate complex academic processes into clear guidance for students and researchers.
Research smarter with Acade
One Academic Agent for literature search, research design, writing, and more.
Try Start free →No credit card required.
