Literature discovery & systematic review — search strategy kit

A practical, step-by-step guide with ready examples, templates, and checklists to find, screen, extract, and synthesize evidence efficiently and rigorously. Includes boolean examples, screening and extraction templates, bias-assessment prompts, bibliometric tips, and reproducibility recommendations.

Quick welcome

If your hunger is to find the best available evidence quickly and reliably, this kit helps you run reproducible literature searches and systematic evidence scans without reinventing the process. It focuses on practical steps you can use for rapid reviews, systematic reviews, and evidence scans—plus templates and examples you can copy into your team workflow.

Why this matters

Poorly scoped, undocumented, or biased searches waste time and lead to missed prior work. The approach below helps you frame an answerable question, search widely, screen reliably, extract consistently, and assess bias so your evidence synthesis is defensible and useful.

Core steps at a glance

  1. Define the question, scope, and inclusion/exclusion criteria.
  2. Build search strings (keywords, synonyms, controlled terms) and adapt them to each database.
  3. Run searches and document everything (dates, databases, filters, exact strings).
  4. Manage citations and deduplicate.
  5. Screen titles/abstracts and full texts (ideally dual screening) and record reasons.
  6. Extract data using a structured template.
  7. Assess risk of bias / study quality.
  8. Synthesize (narrative or quantitative) and report results with reproducible artifacts.

1. Define scope and inclusion/exclusion criteria

Start by making your question explicit. Common frameworks that help are PICO (Population, Intervention, Comparator, Outcome) for interventions or PCC (Population, Concept, Context) for broader scoping. For each element, list:

  • Exact definitions (age ranges, languages, years, study designs allowed).
  • Inclusion criteria (e.g., randomized trials, observational studies, peer-reviewed articles, preprints allowed).
  • Exclusion criteria (e.g., case reports, opinion pieces, non-human studies).

2. Construct search terms and boolean strategy

Combine controlled vocabulary (MeSH, Emtree) where available with free-text synonyms and common spelling variants. Use parentheses, boolean operators, and proximity operators when supported.

Example pattern (clinical topic):

("DiseaseName"[Mesh] OR diseaseName OR "disease name" OR disease*) AND (intervention OR "treatment name" OR "drug name")

Example (technology / process topic):

("machine learning" OR "deep learning" OR "neural network*") AND ("diagnosis" OR "classification" OR "prediction")

Tips:

  • Test search sensitivity by seeing if known key papers are returned.
  • Document every final search string exactly for each database and the date run.
  • Keep a short list of required terms and an expanded synonym list to adapt per database.

3. Databases, grey literature, and search sources

Use a mix of bibliographic databases and domain sources. Typical choices:

  • Biomedical: PubMed / MEDLINE, Embase, Cochrane Library
  • Cross-disciplinary: Scopus, Web of Science
  • Preprints & registries: medRxiv, bioRxiv, arXiv, clinical trial registries
  • Grey literature: government reports, theses, conference proceedings, organizational websites
  • Broad search: Google Scholar (useful but harder to document reproducibly)

4. Run searches and preserve reproducibility

  • Record: database name, full query, date run, filters applied, number of hits.
  • Export results in a consistent format (RIS, BibTeX, EndNote XML) for import into citation managers.
  • Keep a search log (spreadsheet or saved text file) with each saved query.

5. Citation management and deduplication

Import all results into a single citation manager (Zotero, EndNote, Mendeley) or a screening tool (Rayyan, Covidence). Then:

  • Run deduplication and inspect duplicates manually to avoid losing unique records.
  • Keep a copy of the raw exports in case you need to re-run deduplication with different settings.

6. Screening workflow and extraction template

A robust workflow improves reliability and speeds the process.

  • Phases: title/abstract screening → full-text screening → data extraction.
  • Prefer dual independent screening with a third reviewer resolving disagreements.
  • Record explicit reasons for exclusion at full-text stage (this fuels the PRISMA flow diagram).
  • Extraction template (minimum fields): citation, study design, population/sample, intervention/exposure, comparator, outcomes measured, results (numeric), timepoints, funding/conflicts, notes on limitations.

7. Bias assessment checklist

Choose a risk-of-bias tool appropriate to study designs (examples: Cochrane RoB 2 for RCTs, ROBINS-I for non-randomized studies, Newcastle–Ottawa Scale for cohort/case-control, QUADAS-2 for diagnostic studies). If you cannot use a formal tool, at least document:

  • Selection bias risks
  • Confounding and comparability
  • Measurement (outcome/exposure) bias
  • Missing data and selective reporting
  • Funding or conflicts that may influence results

8. Synthesis and reporting

Decide upfront whether a quantitative synthesis (meta-analysis) is plausible. Key outputs to prepare:

  • PRISMA flow diagram (records identified, screened, excluded, included).
  • Study characteristics table (extraction results).
  • Risk-of-bias summary table.
  • Narrative synthesis explaining heterogeneity and patterns; forest plots if doing meta-analysis.
  • GRADE table or equivalent to summarize certainty of evidence if appropriate.

9. Bibliometrics and citation mapping (optional but useful)

To understand influential papers, clusters, or disciplinary spread, consider bibliometric approaches and tools such as VOSviewer, CitNetExplorer, Dimensions or citation network exports. Use co-citation and keyword clustering to find related literatures you might have missed.

10. Reproducibility, preregistration, and automation

  • Preregister protocols where possible (e.g., PROSPERO for clinical reviews) or save a public protocol file.
  • Share full search strings and export files as supplemental data so others can reproduce your search.
  • Use automation carefully: de-duplication, text-mining assisted screening, and machine-learning triage can speed work but always keep human oversight and record algorithm settings and thresholds.

11. Common mistakes and quick fixes

  • Too narrow keywords → broaden with synonyms and wildcards.
  • Over-reliance on a single database → add 2–3 complementary databases.
  • Poor documentation → keep a searchable log of queries and exports.
  • Skipping dual screening → add spot checks or partial dual screening to estimate error rates.

Practical templates (copy-and-adapt)

Screening log columns: RecordID, Title, Authors, Year, Source, Included_TA (Y/N), Included_FT (Y/N), ExclusionReason_FT, Reviewer1, Reviewer2, ConflictResolvedBy.

Extraction table columns: RecordID, Citation, Design, Population, N, Intervention/Exposure, Comparator, Outcomes (names), Effect size / results, Timepoint, Notes on bias, Funding/COI.

Next steps

Use these templates as a starting point. If you want this kit to become an interactive workflow, consider adding a screening form that stores reviewer decisions, auto-builds a PRISMA flow, and exports the extraction table—this makes repeated or living reviews far easier to manage.

Image search phrase: systematic review search strategy


Discussion

Comments and conversation will live here.