The Literature Search Process for a Review Article 0% read

The Literature Search Process for a Review Article

A literature search for a review article should translate the defined review question into essential concept groups, expand each group with relevant search terms, search databases chosen for appropriate coverage, combine the concepts with suitable Boolean and database-specific syntax, test and refine retrieval, then document and verify the final strategy. The aim is not a fixed number of results but a reproducible search that is sufficiently comprehensive for the review’s stated scope.

Because the required breadth varies by review type, discipline, purpose, and scope, no single database, search string, or fixed source count proves completeness; coverage instead depends on source fit and a documented, testable search history. Literature searching retrieves the evidence base, while study selection, evidence organisation, synthesis, and interpretation come later.

Table of Contents

Translate the Review Question Into Search Concepts

An existing review question should be translated into a small set of search concepts that preserve its essential meaning for retrieval across databases.

Searching the full question literally can be too restrictive or inconsistent because some question wording represents alternative terms, optional modifiers, or details better assessed later.

The planning sequence moves from the meaning and scope of the question to distinct concept groups and then prepares those groups for search terms such as keywords and synonyms.

This stage does not form or materially change the review question; it translates an already-formed question into search-ready units.

The number of concept groups depends on which ideas are essential to retrieval for the specific question and review design rather than on a fixed concept count.

The ordered sequence preserves the question's essential meaning while separating concepts needed for retrieval from wording that belongs in term expansion or later screening.

From review question to search concepts

  1. Identify the essential ideas: determine which ideas must be represented for retrieved studies to address the central scope of the review question.
  2. Group each essential idea as one concept: treat wording that represents the same underlying idea as belonging to the same concept group rather than creating separate concept blocks.
  3. Separate nonessential modifiers: distinguish wording that does not need its own search concept because it can be represented through alternative search terms or assessed later during screening.
  4. Prepare each concept for term expansion: use each concept group as the basis for developing relevant keywords, synonyms, and other search terms for the databases being searched.

For example, a review question about whether workplace exercise programs reduce lower-back pain among office workers could treat workplace exercise, lower-back pain, and office workers as candidate search concepts, while alternative wording for those ideas belongs in term expansion.

A narrowly stated study characteristic may instead remain outside the concept groups when it is more appropriately assessed during screening.

Identifying the essential concepts first provides the search-ready units needed for subsequent term expansion.

The image illustrates the translation from the meaning of a review question into search-ready concept groups, complementing the ordered planning sequence without replacing it.

Review question translated into search-ready concept groups for literature searching

Identify the Main Concepts That Must Be Searched

Only the main concepts necessary to preserve the meaning and scope of the review question should become independent search blocks.

A search concept represents an idea needed for relevant retrieval, while wording that is optional or better assessed during screening does not automatically require its own concept block.

The annotated example distinguishes meaning-bearing searchable concepts from a nonessential modifier within one review question.

Review question annotated to distinguish essential search concepts from an optional modifier

Depending on the specific review question, an essential concept may represent a population or subject, phenomenon, intervention or exposure, outcome, or contextual constraint; no single component set applies to every question.

For example, in a question about workplace exercise and lower-back pain among office workers, lower-back pain may be an essential concept because removing it changes the target outcome, whereas a narrowly stated study characteristic may be an optional modifier better handled during screening.

This distinction keeps necessary searchable concepts explicit without allowing nonessential wording to suppress potentially relevant retrieval.

Each approved concept should be expanded with search terms that represent the alternative language authors and indexers may use for the same idea.

Terminology variation affects retrieval because relevant records may express a concept with different keywords, phrases, acronyms, or spelling variants.

Free-text terms can be discovered by examining titles and abstracts in relevant seed papers, while subject headings and other controlled vocabulary should be checked in the thesaurus of each database.

A term-expansion table keeps each alternative expression attached to its parent concept rather than treating the expressions as separate topics.

Free-text keywords represent language that may appear in titles or abstracts, synonyms and spelling variants capture equivalent author terminology, and controlled-vocabulary candidates identify terms to verify against the database's own indexing system.

This structure helps organise term variants without broadening the concept beyond its intended meaning.

Concept Free-text keywords Synonyms or spelling variants Controlled-vocabulary candidates
Stroke stroke cerebrovascular accident, CVA Verify the relevant subject heading in each database thesaurus

For the concept stroke, the table illustrates how a common free-text term, an alternative phrase, and an acronym can remain grouped under one meaning while a database-specific controlled-vocabulary term is verified separately.

True alternatives can belong to the same concept group because they represent the intended idea, whereas a merely related condition should not be treated as a synonym when it broadens that meaning.

Each keyword, synonym, spelling variant, acronym, or verified subject heading should therefore remain mapped to one parent concept when coherent search-term groups are built.

The image illustrates the complementary roles of author terminology and database indexing language within a single concept without reproducing the table's term list.

One search concept mapped to free-text terms and controlled-vocabulary candidates

Choose Databases and Search Platforms for the Review Scope

Database selection should follow the review question, disciplinary scope, evidence types, and required coverage rather than source popularity.

Suitable databases and search platforms are those whose subject coverage and indexed document types fit the evidence the review needs.

Search functionality also matters because the source must support an appropriate way to retrieve records within that scope.

Source suitability depends on what a database covers, how it indexes that material, and which search characteristics it supports rather than on name recognition.

The illustration shows why different search sources can cover overlapping but nonidentical evidence spaces, while the criteria identify the main dimensions for judging source fit.

Review question connected to search sources with overlapping subject and document coverage

For example, a review centred on biomedical journal literature may include PubMed when its disciplinary and document coverage fits the review question, whereas an interdisciplinary review may require sources covering different disciplines or evidence types.

Google Scholar can serve a different search-source role, but broad discoverability by itself does not establish that it fits a particular review scope.

Database choice should therefore remain conditional on the topic and evidence needs, with actual source coverage checked before the final source set is established.

Match Source Coverage to the Review Topic and Question

Source fit is tested by comparing the evidence that must be discoverable for the review question with the subject coverage, journal coverage, document types, and indexing that a source actually provides.

Its search capabilities also affect suitability because searchable fields, controlled vocabulary, and other retrieval functions determine how indexed evidence can be exposed.

The comparison image illustrates how discipline-specific depth and multidisciplinary breadth can produce overlapping but different coverage profiles.

Discipline-specific and multidisciplinary source coverage profiles compared against a review topic

These criteria test discoverability rather than source prestige or result count, and their importance depends on the evidence required by the review.

For example, a discipline-specific source may provide greater depth within a specialised field, while a multidisciplinary source may provide broader cross-disciplinary coverage; either may miss evidence represented more strongly by the other profile.

Suitability therefore depends on whether the source's indexed coverage and search capabilities match the review topic and question, not on breadth or familiarity alone.

Use Multiple Databases When One Source Cannot Cover the Review Scope

Multiple databases are justified when one source cannot adequately represent the disciplines, publication types, or indexing needs within the review scope.

Useful overlap is expected when complementary sources index some of the same literature, but another database adds value when it addresses a plausible coverage gap rather than merely increasing the database count.

The conditions below indicate when additional sources may contribute meaningful coverage.

Overlap between databases does not make the additional source redundant if each may retrieve unique records relevant to the question.

For example, two sources may index many of the same journals while one also indexes relevant publication types or disciplinary material absent from the other, so their combination can extend coverage despite substantial duplication.

The appropriate database mix therefore depends on whether each additional source addresses a credible coverage gap within the review scope, not on reaching a prescribed number of databases.

Build the Search Strategy With Boolean Logic and Database Syntax

A search strategy converts approved concept groups into explicit retrieval logic before the expression is translated into database syntax.

Alternative terms that represent the same concept are normally grouped together first, and separate concept groups are then combined to construct the initial search expression.

This keeps the review concepts visible while making the Boolean logic explicit.

The sequence below constructs the initial search expression in a controlled order so that grouping remains clear before database-specific translation.

Build the Boolean structure

  1. Group alternative terms within each concept: place synonyms, spelling variants, acronyms, and other equivalent terms for one concept together, normally using OR to represent alternatives.
  2. Combine separate concept groups: connect distinct concept groups, normally using AND, so the search expression represents the required intersection between concepts.
  3. Check the grouping: use parentheses or equivalent grouping logic to keep alternative terms attached to the correct concept and prevent unintended changes in meaning.
  4. Prepare for database-specific translation: keep the retrieval logic stable while allowing fields, phrase handling, operators, and other database syntax to be adapted for each search platform.

For example, an initial search expression might be (workplace exercise OR workplace training) AND (lower-back pain OR low back pain).

The expression demonstrates the distinction between grouping alternative terms within one concept and combining separate concepts through Boolean logic.

It is a logical starting point rather than a portable query, because the database syntax may need to change before the search string is used in a specific database.

Combine Concept Groups and Alternative Terms With Boolean Operators

Boolean operators define the logical relationship between search terms and concept groups, which changes what a database may retrieve.

OR combines alternative terms for the same concept, AND requires separate concept groups to occur within the search logic, and NOT excludes specified terms.

The table compares these roles so their typical retrieval effects and main risks are explicit.

Operator Relationship Typical retrieval effect Main risk
OR Combines synonyms or alternative terms representing the same concept May broaden retrieval by allowing records containing any grouped alternative Semantically loose alternatives can introduce records that do not represent the intended concept
AND Requires separate concept groups to intersect May narrow retrieval to records representing the combined concepts Adding unnecessary concept groups can exclude otherwise relevant records
NOT Excludes records containing the specified term or expression May reduce retrieval by removing an unwanted concept or meaning Broad exclusion can also remove relevant records that contain the excluded term in another context

For example, (workplace exercise OR workplace training) AND lower-back pain retrieves the lower-back-pain concept together with either alternative expression for workplace exercise, whereas replacing AND with OR would allow records representing either concept rather than their intersection.

A NOT condition would deliberately exclude specified records, but it should be used cautiously because relevant records may contain the excluded wording for reasons unrelated to the intended exclusion.

Operator choice should therefore reflect the semantic relationship required between concept groups, while exact grouping and database syntax may vary by search interface.

Adapt Search Syntax and Controlled Vocabulary to Each Database

The conceptual search strategy should remain stable while database syntax and indexing language are translated for each database.

Fields, subject headings, phrase searching, truncation, proximity functions, and automatic mapping can differ across search interfaces, so a search expression should not be copied unchanged between platforms.

The table shows how one stable search concept can be translated through database-specific features without changing its semantic intent.

Search feature Database-specific expression Implementation implication
Free-text Apply the concept's keywords and alternative terms in the searchable text fields supported by the database Verify which fields, such as title or abstract, are available and how field tags are written
Controlled vocabulary Map the same concept to relevant subject headings when the database provides an indexed vocabulary Verify the database's own controlled vocabulary rather than transferring a subject heading from another index
Phrase searching Apply the platform's supported phrase-searching syntax when a multiword expression needs to remain together Verify how the database interprets quotation marks or other phrase-search commands
Truncation and wildcards Apply the database-supported symbol or command when a justified word stem or spelling pattern is needed Verify the permitted symbols and their effect before translating the expression
Proximity and automatic mapping Apply proximity operators or automatic mapping only where the database supports and interprets them Check how the interface maps terms or constrains word relationships because these behaviours can differ by platform

Free-text and controlled vocabulary are complementary: free-text represents wording that may appear in titles, abstracts, or keywords, while subject headings represent concepts assigned through a database's indexing system.

Field tags, phrase searching, truncation, proximity operators, and automatic mapping should therefore be translated according to the features supported by the database rather than assumed to behave identically.

Controlled vocabulary should supplement rather than automatically replace relevant free-text searching.

For example, a concept such as stroke can retain stroke and appropriate alternative terms as free-text while also being mapped to the relevant subject heading in a database that provides controlled vocabulary.

In another database, the same concept may require different field tags, phrase handling, truncation rules, proximity syntax, or automatic mapping behaviour.

The semantic intent remains stable, but the database-specific expression should be verified against the platform's current search documentation before the translated strategy is used.

Search refinement is iterative: the initial strategy should be run, inspected, revised when justified, and rerun rather than treated as final after the first search.

The purpose of iteration is to evaluate how the search behaves against the review question, not to reach an arbitrary result count.

Retrieved records provide the evidence needed to decide whether search terms, grouping, or other parts of the strategy need refinement.

The cycle below keeps each change traceable and links revisions to observed retrieval behaviour rather than result volume alone.

Refinement cycle

  1. Run: execute the current search strategy in the selected database and retain the search expression used.
  2. Inspect: review retrieved records for relevance, recurring irrelevant material, terminology patterns, and evidence that important material may be missing.
  3. Revise: refine only the search terms or logic that have a clear reason for change, such as adding a relevant synonym or removing an overly broad alternative term.
  4. Rerun and compare: rerun the revised strategy and compare the new retrieval with the earlier results, checking whether relevance or discoverability changed in the intended direction before making another revision.
  5. Record: document the search history, including the version of the strategy, the change made, and the reason for that refinement.

For example, if inspection shows that relevant records repeatedly use a synonym that is absent from one concept group, adding that term and rerunning the search may retrieve additional relevant records that the earlier version missed.

A refinement should be kept because it improves alignment with the review question or makes retrieval behaviour easier to justify, not simply because the total number of results rises or falls.

Search refinement therefore ends with documented, diagnostic checking of retrieved records rather than a fixed target for result count or number of iterations.

Check Whether Known Relevant Studies Are Retrieved

A known relevant study can act as an exemplar for testing whether the current search strategy retrieves material it should reasonably find.

If the exemplar is missing, that result indicates a need to check the strategy rather than proving one specific search failure.

Retrieving several known studies also does not establish that retrieval is complete.

The checks below identify plausible reasons why a known relevant study may be absent and show what each diagnostic check can reveal.

Terminology, indexing, database coverage, and syntax can each affect retrieval, so a missing study should be treated as a signal that requires diagnosis rather than assigned a cause immediately.

If a known relevant study is missing

For example, suppose a known relevant study is missing even though its topic matches the review question.

Inspection may show that its title or indexing uses an alternative term absent from the corresponding concept group, suggesting a terminology gap rather than a problem with database coverage or syntax.

Adding that justified alternative term and rerunning the search tests whether the adjustment restores retrieval of the exemplar without treating that single result as evidence of completeness.

Adjust Search Breadth and Precision Without Changing the Review Question

Search breadth and precision should be balanced against the review question rather than adjusted toward a fixed result target.

An over-narrow search may have lower sensitivity and miss relevant records, whereas an over-broad search may have lower specificity and retrieve many irrelevant records.

These are diagnostic search states: the appropriate retrieval balance depends on the topic, database, and baseline search behaviour.

Adjustable variables include synonyms, field restrictions, phrase searching, proximity syntax where supported, concept requirements, and exclusions.

Adding valid synonyms or removing an unnecessary concept may broaden retrieval and increase sensitivity, while justified field restrictions or phrase and proximity conditions may narrow retrieval and increase specificity.

The diagnostic table links each observed symptom to a plausible cause, an adjustment, and its expected retrieval effect rather than treating any adjustment as a fixed rule.

Retrieval symptom Likely cause Possible adjustment Expected retrieval effect
Relevant records appear to be missing A concept group may omit valid alternative terminology Add justified synonyms or alternative terms to the affected concept group May broaden retrieval and increase sensitivity by allowing additional relevant terminology
Relevant records are excluded unless they contain an unnecessary concept The baseline search may require a concept that is not essential for retrieval Remove the unnecessary concept requirement while preserving the review question May broaden retrieval and recover relevant records that do not express that concept explicitly
Relevant records are missed because terms occur outside restricted fields Field restrictions may be too narrow for the way relevant records are indexed or described Broaden or remove the restrictive field condition where justified May increase sensitivity by allowing the search terms to match in additional searchable fields
Many irrelevant records contain the required words in unrelated contexts The baseline search may allow terms to occur without a sufficiently specific relationship Use justified phrase searching or proximity syntax where the database supports it May increase specificity and reduce irrelevant records by requiring a closer relationship between terms
Relevant records disappear after an exclusion is applied An exclusion may be removing records that contain both relevant and unwanted meanings Review, narrow, or remove the overly aggressive exclusion May increase sensitivity by restoring records removed by the exclusion

For example, if a baseline search misses relevant records because a concept is limited to one field, relaxing that field restriction may broaden retrieval; if the baseline search instead retrieves many irrelevant records because two words occur independently, supported phrase or proximity syntax may narrow retrieval.

The resulting effect should be evaluated in the database because search features and indexing behaviour differ across platforms and topics, rather than narrowing solely to reduce workload or broadening solely to increase result counts.

The review question remains fixed while retrieval variables are adjusted to improve the balance between sensitivity and specificity.

Extend Search Coverage With Citation Searching and Grey Literature

Supplementary searching can extend database searching when relevant evidence is connected through citation relationships or is poorly represented in conventional indexing.

Citation searching and grey-literature searching address different coverage gaps, so they should be used as complementary methods when the review scope, discipline, or evidence type makes those gaps plausible.

The comparison below distinguishes the evidence each method is designed to locate and the main limitation that should shape its use.

Supplementary method Target evidence Main use Limitation
Citation searching Studies connected to a known relevant or seed study through its references or later citations Follows the citation network to locate connected studies that keyword or subject-heading searches may not expose clearly Retrieval depends on the citation relationships and coverage available from the citation source, so it does not independently establish complete coverage
Grey-literature searching Reports, theses, working papers, government or organisational publications, and other material outside conventional journal publishing Extends coverage when relevant evidence types may be weakly indexed or absent from standard bibliographic databases and may help address evidence affected by publication bias Sources can be distributed across different repositories or websites, with variable indexing and search functionality

Citation searching follows relationships between publications, whereas grey literature searching targets evidence that may sit outside traditional indexed publishing channels.

The first method is most useful when connected studies may be discoverable through a citation network; the second may be useful when the review requires reports, theses, or other nontraditional publication types.

Their functions are distinct even when both extend overall coverage.

Whether either supplementary method is needed depends on the review scope, discipline, expected evidence types, indexing patterns, and the likelihood that publication bias or incomplete database coverage could affect retrieval.

A review focused mainly on well-indexed journal literature may need little supplementary searching, while another review may reasonably require citation searching, grey literature, or both.

Database searching remains the baseline strategy, with supplementary methods added only when they address a credible evidence-discovery gap.

Use Backward and Forward Citation Searching to Find Connected Studies

Backward citation searching follows the references used by a relevant seed study, while forward citation searching identifies later works that cite that seed study.

Start with a seed study that is clearly relevant to the review question, then use the two directions to explore connected studies within its citation network.

The comparison below separates the references a paper used from later works that cite it and shows what each direction may reveal.

Citation-searching method Direction What to check What it may reveal
Backward citation searching From the seed study to earlier cited works Review the seed study's reference list and record relevant candidate studies Earlier studies, source papers, or foundational work connected to the seed study
Forward citation searching From the seed study to later works that cite it Use a cited by function in a citation index, such as Google Scholar or another citation-search platform, and record relevant candidates Later studies that build on, apply, discuss, or otherwise cite the seed study

For example, a relevant seed study may cite an earlier paper that established an important method, while a forward search may identify a later study that applied the same approach in a related context.

The earlier paper is found through the reference list, whereas the later study is found through cited by data, and both can be recorded as candidates for further review.

Citation searching does not map every connected study because discovery depends on the chosen seed studies and on the citation data and platform coverage available.

Search Grey Literature When It Fits the Review Scope

Grey literature should be searched when the review scope plausibly includes relevant evidence disseminated outside conventional academic journal publishing.

It may include reports, theses, conference materials, government publications, and documents produced by organisations or institutions.

The criteria below keep grey-literature searching tied to the review question and evidence needs rather than expanding into an unbounded web search.

When grey literature can add coverage

Searching grey literature often requires adapting search terms and source-specific limits because indexing and search interfaces vary across repositories, government websites, and organisations.

Some sources may provide structured metadata or filters, while others rely on simpler search functions, so the search method and search history should be documented clearly.

Accessibility and source stability can also affect whether a record can be located again during later stages of the review.

Grey literature should not be treated as inherently lower or higher quality than indexed literature because peer review, editorial control, and reporting standards vary by source.

Where the review scope is limited to well-indexed journal evidence, grey-literature searching may add little; where relevant reports, theses, conference materials, or government publications are plausible evidence sources, it may materially extend coverage.

The decision should therefore follow the review scope, evidence needs, indexing conditions, and practical accessibility of the sources being considered.

Document and Verify the Search Strategy Before Ending the Search

Ending a literature search requires both reproducible documentation and a reasoned verification of remaining search-specific coverage concerns. The two requirements are linked, but they answer different methodological questions.

Documentation

Record how the search was executed: the databases or platforms, database-specific search strings, relevant fields and limits, adaptations, search dates, and later revisions. This creates traceability and reproducibility, but it does not claim that every possible study was found.

Verification

Assess whether material coverage concerns remain after the planned searching and refinement. Search further when another reasonable method or source could address a plausible gap; when no proportionate search action is likely to resolve it, document the limitation transparently.

Do not end the search solely because a preferred result count, fixed stopping date, or subjective sense of completion has been reached. No practical search can prove absolute completeness.

Record Databases, Search Dates, Search Strings, and Applied Limits

A reproducible search record must capture the exact source and execution conditions used for each search, including the database, search date, full search string, fields, limits, and any syntax adaptations.

Recording only the database name is insufficient because retrieval can change when the expression, fields, filters, interface, or execution date changes.

These details preserve the conditions under which the search was run.

Translated strategies and later revisions should also remain traceable in the search history.

When a search string is adapted for another database, record the database-specific syntax adaptations and any change in fields or limits; when the strategy is revised, document what changed and why.

Saved search histories or exports can provide practical evidence of these changes when the platform makes them available.

Search record fields

Preserving this search history creates an audit trail that supports reproducibility, later verification, and updating without mixing the search record with study eligibility or evidence-synthesis decisions.

Before ending the search, diagnose any remaining coverage gap as a possible search-design, terminology, indexing, source coverage, or deliberate scope-boundary issue rather than assuming that the evidence base itself is incomplete.

Search completion is a reasoned judgment about whether plausible retrieval blind spots have been addressed, not evidence that every possible record has been found.

Each unresolved issue should therefore lead to a specific diagnostic check before another search adjustment is made.

The checklist below tests whether remaining blind spots arise from the search strategy, database coverage, terminology, applied limits, or supplementary searching.

Each item identifies what the unresolved signal may reflect and whether one final adjustment or transparent documentation is justified.

The diagnosis should distinguish an unintended retrieval problem from a deliberate scope boundary before the search is closed.

Coverage-gap check

A search coverage gap is different from a research gap: the former can result from terminology, indexing, source coverage, or retrieval design, whereas a research gap concerns the evaluated evidence base after relevant literature has been retrieved and assessed.

The later task to identify research gaps should therefore not begin from the absence of search results alone.

If a remaining blind spot has a plausible search remedy, make one justified search adjustment; if it reflects a deliberate scope boundary or cannot reasonably be resolved, document the limitation transparently.

Once the remaining coverage concerns have been addressed or documented and the search is judged sufficiently complete for its stated scope, the workflow can move on to select and screen studies.

That transition marks the end of search verification and the start of evaluating which retrieved records belong in the review.