This chapter is the seminar deck S2 — Systematic Literature Review: A Methodology for Scientific Surveys. Chapter 14 described the social process that produces literature; here the course turns the reader into a producer: how to survey a field systematically, with a repeatable method rather than a personal selection. The stakes are direct — the deck closes by announcing the course artefact: a survey on a topic of some relevance in the fields of Intelligent Systems Engineering. This chapter is the method for writing it.
A literature survey (or, just survey) is a study of the available research works and results concerning a given subject. A good survey should:
Notice the verbs: include, distinguish, provide, place, recall. A survey is not a list of papers — it is a study that synthesises, positions and evaluates.
The goal of a survey is set by its targets. Temporally, three are the possible (ideal) spans:
The time horizon of the survey sets its goal, along with many other choices — organisation, languages, relevance, rationale, …
As far as the audience is concerned, three are the possible targets (simplifying a lot):
Roughly speaking: the Related Work section in most scientific papers is typically organised and written t0–a10; surveys [Ciancarini, 1996; Omicini, 2013] are basically conceived as t10–a1000; books [Mézard and Montanari, 2009] can be easily understood as t∞–a∞.
Surveying is hard, and the deck lists why:
→ A sound methodological approach to survey is clearly required. The deck’s motivation is the course’s own: new topics in computer science and artificial intelligence are going to pop up like popcorn in a pan, gain high relevance in the industry in a short timespan — requiring practitioners to learn them fast, and possibly to share newly-acquired knowledge within an organisation. The ability to study a new scientific/technical topic, gain and share up-to-date critical competence, possibly by producing a well-structured survey, is going to become an essential skill for computer scientists and engineers, as well as for intelligent system experts.
The need for a well-founded methodological approach to literature results clearly emerges in medical research, where the notion of meta-analysis gets early relevance [Lau et al., 1992; Oxman et al., 1993]. The notion of systematic literature review (SLR) basically develops in the healthcare domain [Mulrow, 1994]; it gets popular [White and Schmidt, 2005; Nightingale, 2009] and then somehow formalised in more or less a decade in terms of Cochrane Reviews [Higgins and Green, 2008]. The lineage matters: SLR was born where decisions are matters of life and death, and where a biased reading of the evidence is not an academic inconvenience but a public-health failure.
The deck quotes the Cochrane definition [Higgins and Green, 2008]:
Systematic review. A systematic review attempts to identify, appraise, and synthesise all the empirical evidence that meets pre-specified eligibility criteria to answer a given research question. Researchers conducting systematic reviews use explicit methods aimed at minimising bias, in order to produce more reliable findings that can be used to inform decision making.
Three verbs carry the method: identify (find the evidence), appraise (judge it), synthesise (combine it). Two qualifiers carry the discipline: pre-specified eligibility criteria and explicit methods aimed at minimising bias. Everything in the SLR methodology exists to serve those qualifiers.
The types of SLR distinguished in the Cochrane handbook:
The key features of an SLR, again from Cochrane:
Whereas the need for SLR becomes evident first in healthcare research, the same requirements more or less emerge in the scientific literature everywhere. In the software engineering (SE) area, the same notion is developed first [Kitchenham et al., 2009] and put to test [Kitchenham et al., 2010; Kitchenham, 2012]. Since SE is way closer to our theory and practice, the course refers to that notion henceforth.
Common reasons for performing an SLR [Kitchenham et al., 2009]:
Benefits of SLR [Kitchenham et al., 2009]:
Drawbacks of SLR [Kitchenham et al., 2009]:
The deck contrasts the two practices point by point [Kitchenham et al., 2009]:
The summary is: a standard survey is an expert’s narrative built on a chosen framework; an SLR is a documented process whose every decision — protocol, search, criteria, extraction, quality — is stated in advance and left visible in the final artefact.
SLR is conducted in five main steps:
(1) Research objective & questions. The first conceptual step of SLR is to define the SLR goal; after that, the first technical step is to articulate the corresponding research questions. The review questions drive the whole SLR methodology: the search process must identify the primary studies that address the research questions; the data extraction process must extract the data items needed to answer the questions; the data analysis process must synthesise the data in such a way that the questions can be answered.
(2) Search strategy. The aim of a systematic review is to find as many pieces of scientific literature related to the topic as possible, avoiding all possible bias in the search strategy. The search strategy needs to be totally motivated and explained: sources, keywords, criteria for inclusion / exclusion. The search should be documented to be reproducible — e.g. all sources of documents should be named and referred, and the specific search described.
(3) Study selection. Criteria for inclusion / exclusion should be applied in a clearly-documented and reproducible way; they should be based on the research questions, and pre-defined.
(4) Quality assessment criteria. It is critical to assess the quality of the primary sources for the SLR, in order to: provide still more detailed inclusion/exclusion criteria; investigate whether quality differences provide an explanation for differences in study results; weight the importance of individual studies when results are being synthesised; guide the interpretation of findings and determine the strength of inferences; guide recommendations for further research.
(5) Data extraction & synthesis. Selected studies are to be read for data extraction purposes: specific features, or attributes, are to be extracted from the selected works and framed together, in order to answer the research questions and draw the SLR conclusion.
Reading someone else’s SLR in the SE field can surely help you understand how to conduct an SLR by yourself — both “classical” ones, by Brereton [Budgen and Brereton, 2006; Brereton et al., 2007] or Kitchenham [Kitchenham et al., 2009; Kitchenham et al., 2010], and new ones [Calegari et al., 2021] — the latter being the SLR on logic-based technologies for multi-agent systems, directly relevant to this course’s own subject matter.
The conclusion of the deck is a statement of power and a call to action:
A survey is a study, not a list: it includes all relevant literature, distinguishes assessed knowledge from latest proposals, and provides a rationale and a framework. The systematic literature review turns that ambition into a method — born in medicine (Cochrane), imported into software engineering (Kitchenham) — in five steps: research objective & questions, search strategy, study selection, quality assessment criteria, data extraction & synthesis. For the exams: an SLR is defined by its review protocol, its documented and reproducible search strategy, its explicit inclusion/exclusion criteria, and its quality assessment — all aimed at minimising bias.
A study of the available research works and results concerning a given subject. It should include all the relevant literature on the topic (possibly including past/historical works), clearly distinguish well-assessed knowledge (theories, methods, techniques) from latest proposals, provide a rationale over the subject, come with a framework placing the topic within its research field, and possibly recall ongoing work and future challenges.
Time spans: t0 (here and now — hot topic, everything relevant, potential impact); t10 (the last decade — topics at their supposed peak needing a recap); t∞ (from now back to the start — fully developed topics amenable to re-framing, often as a book). Audiences: a10 (specialists — platform for a rush of research), a1000 (learners — PhD students, young researchers, practitioners), a∞ (educated people on Earth — public understanding). Related-work sections are t0-a10; surveys are t10-a1000; books are t∞-a∞.
Because the amount of relevant material is often overwhelming, incoherent and unmanageable; inclusion/exclusion criteria may hugely vary; the material is heterogeneous (form, source, reliability, scope, goal); surveys are scientific literature by themselves, hence subject to reproducibility and refutability [Popper, 2002], and meta-analysis is essential in many areas; and the wide availability of literature makes surveys possible for almost anyone. New topics in CS/AI pop up fast and gain industry relevance quickly, so the ability to study and survey a topic becomes an essential skill.
In medical research, where meta-analysis gets early relevance [Lau et al., 1992; Oxman et al., 1993]; the notion of SLR basically develops in the healthcare domain [Mulrow, 1994], gets popular [White and Schmidt, 2005; Nightingale, 2009] and is formalised in about a decade in terms of Cochrane Reviews [Higgins and Green, 2008].
A systematic review attempts to identify, appraise, and synthesise all the empirical evidence that meets pre-specified eligibility criteria to answer a given research question. Researchers conducting systematic reviews use explicit methods aimed at minimising bias, in order to produce more reliable findings that can be used to inform decision making [Higgins and Green, 2008].
Intervention reviews (benefits and harms of interventions); diagnostic test accuracy reviews (how well a test diagnoses a disease); methodology reviews (how systematic reviews and clinical trials are conducted and reported); qualitative reviews (synthesising qualitative evidence beyond effectiveness); prognosis reviews (probable course or future outcomes); overviews (summarising multiple intervention reviews for a single condition or health problem).
A clearly stated set of objectives with pre-defined eligibility criteria for studies; an explicit, reproducible methodology; a systematic search attempting to identify all studies meeting the eligibility criteria; an assessment of the validity of the findings of the included studies (e.g. risk of bias); a systematic presentation and synthesis of the characteristics and findings of the included studies.
Kitchenham and colleagues [Kitchenham et al., 2009; 2010; 2012]. Common reasons: to summarise the existing evidence concerning a treatment or technology; to identify gaps in current research so as to suggest areas for further investigation; to provide a framework, background, landscape against which to position new research activities.
Benefits: a well-defined methodology makes results less likely to be biased (though it does not protect against biases in primary studies); information about effects across a wide range of settings and empirical methods (consistent results → robust and transferable evidence; otherwise, sources of variation can be studied); meta-analytic techniques increase the likelihood of detecting real effects. Drawbacks: considerably more effort than traditional reviews (time, people, time for publication); results cannot be easily predicted, which may make it difficult to devise an overall rationale.
SLRs start by defining a review protocol (research question + methods); are based on a defined search strategy aiming to detect as much relevant literature as possible; document the search strategy so readers can assess rigour, completeness and repeatability (though digital-library searches are almost impossible to perfectly replicate); require explicit inclusion and exclusion criteria for each potential primary study; and specify the information to be obtained from each study, including quality criteria.
(1) Research objective & questions: define the goal, then articulate the research questions that drive the whole methodology; (2) search strategy: find as much relevant literature as possible, totally motivated and explained, documented to be reproducible (sources, keywords, criteria); (3) study selection: apply pre-defined inclusion/exclusion criteria in a documented, reproducible way, based on the research questions; (4) quality assessment criteria: assess the quality of primary sources to refine criteria, explain differences, weight studies, guide interpretation and recommendations; (5) data extraction & synthesis: extract features/attributes from selected studies and frame them together to answer the research questions and draw the conclusion.
A sound documental artefact for the course, reporting a survey on a topic of some relevance in the fields of Intelligent Systems Engineering — using the SLR methodology to collect literature systematically, summarise and understand scientific and technical results, and make them usable for system engineering, scientific research, and organisation or political decision making.