Abstract
Background
Natural language processing (NLP) techniques offer promising solutions for semi-automating the time-consuming process of abstract screening in systematic reviews. The exponential growth of published literature has created significant bottlenecks, with review teams manually assessing thousands of abstracts over weeks to months. Single reviewers can miss 5-13% of relevant studies, necessitating dual screening that further increases workload. Advances in artificial intelligence, including deep learning models such as BERT and its successors, show potential for automating this critical step, but comprehensive evidence on optimal approaches, performance, and practical feasibility remains limited.
Objectives
This systematic review aimed to assess techniques, performance, and feasibility of NLP approaches for title and abstract screening by characterizing the range of NLP methods used, summarizing performance on key metrics like workload reduction and recall, evaluating real-world implementation feasibility, and identifying research gaps and future directions.
Search Methods
We searched PubMed, Web of Science, Embase, CINAHL, The Cochrane Library, Scopus, and gray literature sources from inception to December 2024. The search strategy, developed with an information specialist and peer-reviewed using PRESS guidelines, targeted keywords related to natural language processing, machine learning, abstract screening, and systematic reviews. Additional sources included conference proceedings, preprint servers, reference lists, forward citation tracking, and expert consultation.
Selection Criteria
We included primary studies of any design describing development or evaluation of NLP techniques for automating title and abstract screening in evidence syntheses. Eligible studies reported on NLP methods, screening performance (workload reduction, recall, precision), or implementation feasibility. Studies using only rule-based approaches without machine learning, systematic reviews of NLP methods, commentaries, and conference abstracts were excluded. No language or date restrictions were applied.
Data Collection and Analysis
Two reviewers independently screened titles, abstracts, and full texts using Covidence software, with disagreements resolved through discussion. Data extraction covered study characteristics, NLP techniques, training approaches, performance metrics, and feasibility considerations. Risk of bias was assessed using a modified ROBIS tool. Given diverse techniques and outcomes, we conducted narrative synthesis following SWiM guidelines, grouping studies by NLP approach.
Main Results
From 4,105 records, 19 studies met inclusion criteria, with 68.4% published since 2023, reflecting rapid field advancement. Studies employed diverse approaches from traditional machine learning (Support Vector Machines, Random Forests) to advanced deep learning models, particularly BERT variants. Most achieved >90% recall with workload reductions of 13-96%, representing substantial time savings. Deep learning models with transfer learning consistently outperformed traditional approaches. However, implementation faced significant barriers including requirements for high-quality training data, specialized computational resources, technical expertise, and user-friendly interfaces. Performance was generally better for targeted reviews with lower inclusion prevalence.
Authors’ Conclusions
NLP techniques, especially deep learning with transfer learning, show substantial promise for semi-automating abstract screening with potential for large workload savings while maintaining high recall. However, challenges remain regarding training data quality, computational requirements, technical expertise needs, and user-centered design. Realizing full potential requires interdisciplinary collaboration to develop reliable, generalizable tools integrating seamlessly with human expertise and existing workflows. Future priorities include creating standardized datasets, conducting prospective evaluations, developing user-friendly interfaces, and establishing implementation best practices to revolutionize evidence synthesis efficiency.
Plain Language Summary
Natural language processing substantially reduces abstract screening workload but requires significant expertise and specialized computing resources.
Natural language processing techniques can reduce manual abstract screening workload by 30-90% while maintaining over 90% recall of relevant studies, with deep learning models showing the greatest promise for systematic review automation.
Problem statement: Systematic reviews are the highest level of evidence for policy and practice, but exponential literature growth has made abstract screening a major bottleneck, taking weeks to months. Single reviewers miss 5-13% of relevant studies, making dual independent screening the gold standard, at significant cost in time and resources.
This systematic review examines the techniques, performance, and feasibility of natural language processing methods for automating title and abstract screening in systematic reviews.
This review includes 19 studies that evaluated NLP techniques for automating abstract screening in systematic reviews, rapid reviews, scoping reviews, and other evidence syntheses. Studies came from diverse global regions, with 68.4% published in 2023 or later. Studies varied in design and corpus size (hundreds to tens of thousands of articles), mostly from medical literature, and demonstrated generally good methodological quality.
NLP techniques consistently demonstrate substantial workload reductions while maintaining high recall of relevant studies. Workload reductions range from 13-96%, with most studies achieving reductions of 30-60%. Most studies achieve recall rates exceeding 90%, meaning they successfully identify over 90% of relevant articles. Precision varied widely (10-99%), primarily affecting the volume of articles requiring manual review rather than the risk of missing relevant studies.
Deep learning models, particularly those leveraging transfer learning with large pretrained language models like BERT and its variants (BioBERT, PubMedBERT), consistently outperform traditional machine learning approaches. Traditional approaches such as Support Vector Machines perform well but are generally outperformed by these modern architectures.
Several key factors influence the practical implementation of NLP systems. High-quality training data and specialist technical expertise are prerequisites. Computational resources vary from standard computing for simpler models to specialized GPU infrastructure for advanced deep learning approaches. User-friendly interfaces and domain generalizability remain key challenges for broader adoption.
NLP offers substantial promise for reducing the abstract screening burden. Workload reductions of 30-90% could accelerate evidence synthesis and policy translation. However, successful implementation requires careful planning, technical expertise, and adequate computational resources. Standardized datasets, user-friendly tools, and clearer best-practice guidance are needed to realise this potential.
The review authors searched for studies up to December 2024.
Keywords
Background
Evidence synthesis through systematic reviews and meta-analyses is considered the highest level of evidence to inform healthcare policy and practice (Murad et al., 2016). However, the exponential expansion of the published literature base (Bornmann & Mutz, 2014) has made producing timely reviews increasingly resource intensive (Bastian et al., 2010). A critical bottleneck is abstract screening, where review teams manually assess thousands of titles and abstracts to identify potentially eligible studies for inclusion. This is a time-consuming, error-prone process that can take weeks to months (Borah et al., 2017). Studies have shown that single human reviewers can miss a substantial proportion of relevant studies, ranging from 5% to over 13% (Gartlehner et al., 2020). Dual independent screening is therefore considered the gold standard (Stoll et al., 2019). However, this further increases the workload and resource requirements of the screening process.
Natural language processing (NLP), a field of artificial intelligence (AI) focused on training computers to understand and manipulate human language, offers promise for semi-automating this step (Marshall & Wallace, 2019). NLP systems can be trained on a subset of reviewer decisions to predict the relevance of remaining records, allowing for prioritization or even automatic exclusion. This has the potential to substantially reduce human workload without sacrificing quality.
Recent years have seen tremendous advances in NLP, fueled by the development of powerful deep learning models that can learn rich linguistic representations from vast text corpora (Devlin et al., 2018). Models like BERT (Bidirectional Encoder Representations from Transformers) (Vaswani et al., 2017), pretrained on millions of documents, can be fine-tuned for custom tasks with minimal labeled data, an approach known as transfer learning. While BERT was introduced in 2018 and is now well-established, it and its domain-specific variants (BioBERT, PubMedBERT, SciBERT) remain the most commonly evaluated models for screening. The field continues to evolve rapidly, with newer architectures emerging at pace. These models have driven impressive performance gains on a range of natural language tasks (Minaee et al., 2020).
In the context of abstract screening, numerous studies have explored applications of NLP, from conventional machine learning approaches like support vector machines and naive Bayes (Cohen, 2008; Przybyła et al., 2018; Rathbone et al., 2015; Wallace et al., 2012), to deep learning models like convolutional neural networks (Przybyła et al., 2018), and most recently, fine-tuning of large pretrained transformer models (Gates et al., 2019). However, a systematic synthesis of these methods, their performance, and feasibility is currently lacking.
More recently, large language models (LLMs) such as GPT-4 have emerged as potential screening tools. Our eligibility criteria did not exclude LLM-based studies; however, none of the identified studies meeting all inclusion criteria employed standalone LLM approaches (e.g., zero-shot prompting) as their primary method. We acknowledge this as a rapidly evolving area and future updates should prioritize this literature.
Prior reviews on the topic have been narrow in scope, non-systematic in their methodology, or are now outdated given the rapid pace of advancement in the field (O'Mara-Eves et al., 2015; Tsafnat et al., 2014). Furthermore, prior reviews have given limited attention to the practical feasibility and resource requirements of implementing NLP screening systems. The recent review by Blaizot et al. (Blaizot et al., 2022) focused on evaluating AI methods used in SR, but did not provide a detailed description of the specific NLP techniques employed. Other recent reviews like those by van Dinter et al. (van Dinter et al., 2021) and Boateng et al. (Ofori-Boateng et al., 2024) have looked at automation of SR more broadly, with NLP being just one component. The systematic review by Sundaram and Berleant (Sundaram & Berleant, 2023), while focused on automating systematic literature reviews with NLP and text mining, did not provide a detailed analysis of the performance and feasibility of different NLP techniques for each stage of the systematic review process. These studies highlight the potential of NLP but do not provide an in-depth examination of the techniques and their comparative performance.
Implementing NLP systems for abstract screening in real-world systematic reviews requires careful consideration of several feasibility factors. These include the availability of suitable training data that is generalizable to the review topic, the computational resources and technical expertise needed to develop and deploy models, and the potential challenges of integrating these systems into existing systematic review workflows and software platforms (Olorisade et al., 2016). In-depth examination of these feasibility considerations in the context of real-world case studies is also currently lacking.
This information is critical for systematic review teams considering using automation tools and for informing future research and development priorities. Therefore, there is a need for an up-to-date, comprehensive review that focuses specifically on NLP methods for abstract screening, their inner workings, performance metrics, and real-world feasibility considerations. This review aims to address that gap.
Objectives
The objective of this systematic review is to comprehensively summarize the evidence on the techniques, performance, and feasibility of NLP methods for title and abstract screening in systematic reviews. Specifically, we aimed to. (1) Characterize the range of NLP techniques used, including feature engineering, modeling architectures, and training approaches. (2) Summarize the performance of systems on key metrics like workload reduction, recall, and precision and explore effect modifiers like training set size and screening prevalence. (3) Evaluate the feasibility of implementing systems in real-world settings, considering factors such as availability and generalizability of training data, computational resource requirements, and the need for specialized NLP and machine learning expertise. (4) Highlight gaps in current methods and promising future directions.
While our primary focus is on healthcare and biomedical domains where systematic reviews are most prevalent, we also sought to include applications across diverse fields including biosciences, social sciences, education, and other disciplines where evidence synthesis methods are increasingly employed. This broader perspective allows for cross-domain insights and identification of transferable techniques and approaches.
Methods
This systematic review followed a pre-registered protocol (PROSPERO CRD42024615153) and is reported in accordance with the PRISMA 2020 statement (Page et al., 2021). A completed PRISMA 2020 checklist is provided in Supplemental Table S1.
Eligibility Criteria
We included primary studies of any design that described the development or evaluation of NLP techniques to partially or fully automate title and abstract screening for a systematic review, rapid review, scoping review or other evidence synthesis. For the purposes of this review, we defined NLP broadly to encompass any computational method that processes or represents natural language text using machine learning, including but not limited to supervised classification (e.g., support vector machines, random forests, logistic regression), deep learning (e.g., convolutional neural networks, recurrent neural networks), transformer-based models (e.g., BERT and its variants), topic modeling with machine learning components, and active learning or human-in-the-loop systems employing these methods. All forms of screening assistance were eligible, including fully automated exclusion, screening prioritization, and decision support tools, provided they incorporated an NLP or machine learning component. To be eligible, studies had to report on at least one of the following.
(a) the NLP methods used (e.g. specific techniques, feature engineering, training data) b) screening prioritization performance (e.g. workload reduction, recall, precision) c) feasibility of use (e.g. usability, resource requirements, generalizability).
We included studies with various comparison approaches, including both traditional control comparisons (e.g., manual screening versus NLP-assisted screening) and active comparisons between different NLP/ML methods or implementations. Both studies comparing NLP to manual screening and studies comparing different NLP methods were eligible. This inclusive approach allows for a more comprehensive assessment of the relative strengths and limitations of different automated screening techniques as they evolve in this rapidly advancing field.
Studies that only used rule-based or string-matching approaches without any machine learning components were excluded. Systematic reviews of NLP methods were excluded but their references were screened. Commentaries, editorials, and conference abstracts were excluded. No language or date restrictions were applied.
Information Sources and Search Strategy
We searched MEDLINE (via PubMed), Web of Science Core Collection (SCI-Expanded, SSCI, ESCI), Embase (via Ovid), CINAHL (via EBSCOhost), The Cochrane Library (via Wiley), and Scopus (via Elsevier) from inception to December 2024. A search strategy was developed in consultation with an information specialist and peer-reviewed using the PRESS checklist (McGowan et al., 2016). The search utilized keywords and, where available, subject headings (e.g., MeSH terms in PubMed) related to “natural language processing”, “machine learning”, “text mining”, “abstract screening”, “systematic reviews”, and “evidence synthesis”. The search string utilized for PubMed is as follows:
((“natural language processing” OR “NLP” OR “machine learning” OR “ML” OR “artificial intelligence” OR “AI” OR “deep learning” OR “neural network*” OR “text mining” OR “automated screening” OR “semi-automated screening” OR “computer-assisted screening” OR “automated review*” OR “transfer learning” OR “BERT” OR “transformer model*” OR “language model*”) AND (“systematic review*” OR “evidence synthesis” OR “systematic literature review” OR “meta-analysis” OR “scoping review” OR “rapid review” OR “systematic map*” OR “evidence map*” OR “umbrella review” OR “overview of reviews”) AND (“abstract screening” OR “title screening” OR “citation screening” OR “study selection” OR “article selection” OR “reference screening” OR “literature screening” OR “screening process” OR “screening phase” OR “screening step” OR “record screening” OR “screening automation” OR “automated screening” OR “screening prioritization” OR “workload reduction” OR “screening efficiency” OR “screening accuracy” OR “screening performance”)).
The search strategy was adapted for each database, taking into account differences in syntax and controlled vocabulary (details in Supplemental Material). In addition to the electronic database search, the reference lists of included studies and relevant review articles was hand-searched to identify additional eligible studies.
We also searched gray literature sources, including Google Scholar, conference proceedings (using the Web of Science Conference Proceedings Citation Index), dissertations and theses (using ProQuest), and preprint servers (medRxiv, bioRxiv, arXiv, and Open Science Framework Preprints). Finally, we reviewed the reference lists of included studies and conducted forward citation tracking in Google Scholar. Content experts were contacted to identify additional studies.
Complete search strategies for all databases, including grey literature sources, dates, and hit counts, are reported in Supplemental Appendix S1, following PRISMA-S recommendations (Rethlefsen et al., 2021).
Prior to finalizing this review, we conducted searches in PROSPERO and the Open Science Framework (OSF) to identify any ongoing or recently completed systematic reviews on this topic to avoid duplication and ensure our review provides unique contributions to the literature.
Screening and Selection Process
Search results were deduplicated and imported into Covidence systematic review software (Veritas Health Innovation, Melbourne, Australia) (Innovation, 2025). Two reviewers (RS, ZG) independently screened titles and abstracts, with disagreements resolved through discussion. The full texts of records included at the title/abstract stage were then independently assessed by the same two reviewers, again with disagreements resolved through discussion or consultation with a third reviewer (AM) where necessary. Reasons for full text exclusion were recorded. A flow diagram was prepared summarizing the study selection process.
Data Extraction
A standardized form was developed in Excel to extract key data items, including bibliographic information, study design, evidence synthesis and corpus characteristics, NLP techniques and architectures, training and validation approaches, performance outcomes (e.g. workload reduction, recall, precision, F1), computational resource requirements, and generalizability and feasibility considerations. The form was piloted on 5 studies. Two reviewers (RS, ZG) independently extracted data from each study. Discrepancies were resolved through discussion.
Quality Assessment
The risk of bias of each included study was assessed using an adapted ROBIS tool (Whiting et al., 2016). Domain 1 assessed whether datasets were relevant and clearly defined; Domain 2 assessed description of data sources and selection; Domain 3 assessed completeness of reporting on feature engineering, model parameters, and validation; Domain 4 assessed appropriateness of analytic methods and results reporting. The adapted tool is provided in Supplemental Table S2. Two reviewers (RS, ZG) independently applied the tool, with any discrepancies resolved by discussion. Studies were rated as low, high, or unclear risk of bias across four domains: 1) relevance of datasets to the review question, 2) identification and selection of studies into the dataset, 3) data extraction and study appraisal, and 4) synthesis and findings.
Data Synthesis
Given the diversity of techniques and outcome measures, we conducted a narrative synthesis in accordance with the Synthesis Without Meta-analysis (SWiM) guideline (Campbell et al., 2020). Studies were grouped by the primary NLP technique used (e.g. support vector machines, convolutional neural networks, transformer models). Within each group, we narratively characterized the feature engineering approaches, training and validation techniques, performance outcomes, and feasibility considerations.
Results
Study Selection
The searches retrieved 4,105 records from bibliographic databases and 80 additional records from other sources. After removing 2,948 references (2,848 duplicates identified by Covidence, 56 duplicates identified manually, and 44 for other reasons), 1,237 records underwent title and abstract screening. Of these, 831 were excluded at the screening stage, leaving 406 records for full-text retrieval. After 35 records could not be retrieved, 371 full texts were assessed for eligibility. A total of 352 studies were excluded at the full-text stage, with the most common reasons being: not utilizing NLP techniques (n = 156), not evaluating screening performance (n = 82), not being a primary study (n = 54), and being conference abstracts only (n = 32). A detailed table of full-text exclusions with categorized reasons and frequencies is provided in Supplemental Table S3. Additional reasons for exclusion included being review articles without original data (n = 18), full text not being available (n = 6), and non-English language (n = 4). Ultimately, 19 studies met all inclusion criteria and were included in the review (Bravo et al., 2021; Campos et al., 2024; Chai et al., 2021; Du et al., 2024; Gartlehner et al., 2019; Hamel et al., 2020; Hasny et al., 2023; Kebede et al., 2023; Manion et al., 2024; Masoumi et al., 2024; Moreno-Garcia et al., 2023; Natukunda & Muchene, 2023; Ng et al., 2023; Orel et al., 2023; Perlman-Arrow et al., 2023; Pham et al., 2021; Pilz et al., 2024; Qin et al., 2021; Wong et al., 2024). A PRISMA flow diagram detailing the study selection process is shown in Figure 1. PRISMA Flow Diagram
Study Characteristics and Corpus Details
Study and Corpus Characteristics
Legends: SLR: Systematic Literature Review; HPV: Human Papillomavirus; PAPD: Physical Activity and Parkinson’s Disease; PPU: Perforated Peptic Ulcer; INTRISSI: Intensive Care Unit Related Studies; HIV: Human Immunodeficiency Virus; SARS-CoV-2: Severe Acute Respiratory Syndrome Coronavirus 2; NLP: Natural Language Processing.
The study designs employed in these investigations varied, with methodological studies, model development and validation studies, and experimental or comparative evaluation studies being the most common. This diversity in study designs reflects the multifaceted nature of research in this area, ranging from the development of novel NLP techniques to the evaluation of their performance and feasibility in real-world scenarios.
The corpora used in these studies were primarily sourced from medical and health science databases, such as PubMed, Embase, and Cochrane Library. This predominance of medical and health-related literature aligns with the traditional domains where systematic reviews and evidence synthesis are most widely conducted and where the need for efficient abstract screening is particularly pronounced. However, a few studies also explored the application of NLP techniques in other fields, such as education and psychology (Campos et al., 2024) and space medicine (Hasny et al., 2023), demonstrating the potential for cross-domain transfer of these approaches.
The size of the corpora varied considerably across studies, ranging from a few hundred to tens of thousands of articles. For instance, Chai et al. (Chai et al., 2021) worked with a relatively small corpus of 306 articles, while Hamel et al. (Hamel et al., 2020) utilized much larger datasets comprising 69,663 articles. This wide range in corpus sizes highlights the scalability of NLP techniques and their ability to handle both small-scale and large-scale screening tasks.
The included/excluded ratios, which indicate the proportion of relevant articles within the screened corpora, also exhibited substantial variability. The percentage of included articles ranged from as low as 2.2% (Moreno-Garcia et al., 2023) to as high as 31.7% (Du et al., 2024). This variability can be attributed to factors such as the breadth and specificity of the research questions, the search strategies employed, and the inherent characteristics of the target domains. Studies with lower inclusion rates (e.g. (Masoumi et al., 2024; Orel et al., 2023)) typically dealt with more focused research questions or utilized more stringent eligibility criteria, resulting in a smaller proportion of relevant articles within the screened corpus.
NLP Methods and Technical Implementation
NLP Methods and Technical Implementation
Legends: LDA: Latent Dirichlet Allocation; RF: Random Forest; BERT: Bidirectional Encoder Representations from Transformers; SVD: Singular Value Decomposition; LightGBM: Light Gradient Boosting Machine; SVM: Support Vector Machine; LR: Logistic Regression; TF-IDF: Term Frequency-Inverse Document Frequency; BioBERT: Bidirectional Encoder Representations from Transformers for Biomedical Text Mining; SciBERT: Scientific BERT; SBERT: Sentence BERT; PaCMAP: Pairwise Controlled Manifold Approximation; HDBSCAN: Hierarchical Density-Based Spatial Clustering of Applications with Noise; k-NN: k-Nearest Neighbors; CRF: Conditional Random Fields; LSTM: Long Short-Term Memory; BoW: Bag of Words; AWS: Amazon Web Services; ML: Machine Learning; NLP: Natural Language Processing; NB: Naive Bayes.
Several studies in this review leveraged BERT and its domain-specific variants (e.g., BioBERT, PubMedBERT) for abstract screening (Bravo et al., 2021; Chai et al., 2021; Hasny et al., 2023; Manion et al., 2024; Ng et al., 2023; Perlman-Arrow et al., 2023; Qin et al., 2021; Wong et al., 2024). These pretrained language models have the advantage of capturing rich semantic information and contextual relationships within the text, leading to improved performance compared to traditional machine learning approaches (Devlin et al., 2018; Lee et al., 2020).
Topic modeling techniques, such as Latent Dirichlet Allocation (LDA), were also employed in some studies (Natukunda & Muchene, 2023). Topic modeling aims to discover latent themes or topics within a corpus of documents and can be useful for identifying relevant articles based on their thematic similarities (Blei et al., 2003).
Feature engineering, which involves extracting and representing meaningful features from the text data, played a crucial role in the performance of the NLP models. The most common feature representation techniques included term frequency-inverse document frequency (TF-IDF) matrices (Du et al., 2024; Kebede et al., 2023; Pilz et al., 2024), word embeddings (Pham et al., 2021), and transformer-based embeddings (Bravo et al., 2021; Manion et al., 2024; Ng et al., 2023; Qin et al., 2021). These representations capture different aspects of the text, such as word frequency, semantic relationships, and contextual information, and serve as input to the machine learning or deep learning models (Mikolov et al., 2013; Pennington et al., 2014).
The training and validation strategies employed in the studies varied depending on the specific objectives and the nature of the datasets. Some studies used temporal splits, where the model was trained on articles published before a certain date and validated on articles published after that date (Qin et al., 2021). Others employed cross-validation techniques, such as k-fold cross-validation (Moreno-Garcia et al., 2023), to assess the model’s performance on different subsets of the data. Hold-out validation, where a portion of the dataset is reserved for testing, was also commonly used (Chai et al., 2021; Hasny et al., 2023; Wong et al., 2024).
The majority of the studies utilized Python programming language and its associated libraries, such as scikit-learn, TensorFlow, and PyTorch, for implementing the NLP models. Python has a rich ecosystem of NLP and machine learning tools, making it a popular choice for researchers and practitioners in this field (Pedregosa et al., 2011). Some studies also used R programming language (Kebede et al., 2023; Natukunda & Muchene, 2023; Pilz et al., 2024) or specialized software such as DistillerSR (Gartlehner et al., 2019; Hamel et al., 2020) for their analyses.
Performance Metrics and Implementation Feasibility
Performance Metrics and Implementation Feasibility
Legends: WSS@95: Work Saved over Sampling at 95% recall; AUC: Area Under the Curve; GPU: Graphics Processing Unit; RAM: Random Access Memory; NNR: Number Needed to Read; F1: F1 Score (harmonic mean of precision and recall); CDSMOTE: Class-Decomposition SMOTE (Synthetic Minority Over-sampling Technique); OPOT: One Person One Tool; PIO: Population, Intervention, Outcome; ASReview: Active learning for Systematic Reviews; UI: User Interface; ML: Machine Learning; AWS: Amazon Web Services; NLP: Natural Language Processing.
Precision, which represents the proportion of articles identified as relevant by the model that are actually relevant, was another commonly reported metric. High precision values indicate a lower false positive rate and reduce the manual effort required to review irrelevant articles. The precision values varied widely across studies, ranging from 10% to 99% (Moreno-Garcia et al., 2023), depending on the specific datasets and models used.
The F1-score, which is the harmonic mean of recall and precision, provides a balanced measure of the model’s overall performance. Studies reporting F1-scores demonstrated values ranging from 0.15 to 0.99 (Moreno-Garcia et al., 2023), with some achieving scores of over 0.90 (Perlman-Arrow et al., 2023).
Workload reduction, which quantifies the percentage of articles that can be safely excluded from manual review without compromising the comprehensiveness of the evidence synthesis, is a pragmatic metric that directly reflects the potential efficiency gains offered by NLP-assisted screening. The workload reduction values reported in the included studies ranged from 31% (Wong et al., 2024) to 96% (Chai et al., 2021), with several studies demonstrating substantial reductions of over 50% (Hasny et al., 2023; Kebede et al., 2023; Orel et al., 2023; Perlman-Arrow et al., 2023).
Computational requirements ranged from standard desktop computing to GPU-accelerated and cloud-based infrastructure. Implementation details are reported in Table 3, with external validation and cross-domain testing details provided in Supplemental Table S4.
In summary, this systematic review highlights the significant progress made in the application of NLP techniques for semi-automating the abstract screening process in evidence synthesis. The included studies demonstrate the potential of various NLP approaches, ranging from traditional machine learning algorithms to state-of-the-art deep learning models, in reducing the workload and time required for manual screening while maintaining high levels of recall and precision. However, the generalizability and practical feasibility of these techniques are influenced by factors such as the availability of high-quality training data, domain expertise, computational resources, and user-friendliness. Further research and development efforts are needed to address these challenges and facilitate the widespread adoption of NLP-assisted abstract screening in real-world evidence synthesis projects.
Risk of Bias Assessment
Figure 2 below shows the results of the Risk of Bias analysis, indicating that the majority of studies maintained good methodological quality with particular strengths in study eligibility criteria (Domain 1) and analysis methods (Domain 4). Domain 3 (Data Collection and Study Appraisal) showed the most variability across studies. ROBIS Risk of Analysis
The assessment revealed that the majority of studies demonstrated low risk of bias across domains, particularly in study eligibility criteria (Domain 1: 15/19 studies, 78.9%) and analysis and synthesis methods (Domain 4: 14/19 studies, 73.7%). Data collection and study appraisal (Domain 3) showed the highest proportion of unclear risk assessments (11/19 studies, 57.9%), primarily due to limited reporting of validation procedures and incomplete documentation of data extraction processes (Figure 3). ROBIS Risk of Bias bar Graph
Notable high-quality studies with low risk of bias across all domains included Pham et al., 2021, Du et al., 2024, Gartlehner et al., 2019, Hamel et al., 2020, Perlman-Arrow et al., 2023, Moreno-Garcia et al., 2023, Pilz et al., 2024, and Campos et al., 2024. These studies demonstrated comprehensive search strategies, clear selection processes, standardized data collection methods, and robust analysis approaches. Studies with unclear risk assessments typically showed limitations in documenting validation procedures or providing complete details of their selection processes.
No studies were assessed as having high risk of bias in any domain, suggesting an overall good methodological quality across the included literature. However, improvements in reporting data collection procedures and validation methods could strengthen future studies in this field.
The assessment emphasizes the need for detailed reporting of data collection and validation procedures, clear documentation of study selection processes, standardized quality assessment approaches for automated screening studies, and comprehensive documentation of methodological decisions.
Discussion
This systematic review provides a comprehensive synthesis of the rapidly evolving evidence base on applications of NLP techniques to support title and abstract screening in systematic reviews. The 19 included studies evaluated a wide range of NLP approaches, from conventional machine learning algorithms to cutting-edge deep learning architectures and large pretrained transformer language models.
The evidence suggests that NLP-assisted screening can yield substantial workload reductions, often in the range of 30-60%, and sometimes up to 90% or more, while maintaining high recall of relevant studies. This translates into considerable time savings, from several hours to multiple weeks of person-time, depending on the size of the review. Performance tended to be better for reviews with a more targeted scope and lower inclusion prevalence. State-of-the-art deep learning models, particularly those leveraging transfer learning with large pretrained language models like BERT, consistently outperformed traditional machine learning approaches across various datasets and research domains.
However, the reliability of these tools depends heavily on the quality and quantity of available training data. Most studies to date have relied on “convenient” datasets derived from previously completed reviews, and the generalizability of performance to new review topics is uncertain. Even with advanced techniques, achieving stable results required screening hundreds to thousands of records to fine-tune models, which may still pose a burden for smaller, rapid reviews or emerging research areas where relevant training data is scarce. Furthermore, several studies highlighted potential risks of bias or limited applicability when training datasets were not representative of the target body of literature, particularly for identifying research from low-resource settings or marginalized populations. This underscores the critical importance of thoughtful training data curation.
From an implementation standpoint, while the most advanced deep learning models boast impressive accuracy, they are computationally demanding and typically necessitate specialized technical expertise to develop, validate, and deploy. The “black box” nature of these complex models can also pose challenges for user trust, interpretability, and acceptance. To date, there has been limited emphasis on prospective, real-world evaluation of user-centered screening tools that integrate seamlessly with existing systematic review management software and workflows.
The computational resources required for implementing the NLP models varied depending on the complexity of the algorithms and the size of the datasets. While some studies utilized standard computing resources (Orel et al., 2023), others leveraged high-performance computing infrastructure, such as GPU acceleration (Hasny et al., 2023; Qin et al., 2021) or cloud-based platforms (Manion et al., 2024; Perlman-Arrow et al., 2023). The use of specialized hardware and distributed computing resources can significantly reduce the computational time and enable the processing of large-scale datasets (LeCun et al., 2015).
The generalizability of the NLP models is an important consideration for their practical application across different domains and contexts. While some studies focused on specific medical domains (Masoumi et al., 2024; Wong et al., 2024), others demonstrated the cross-domain adaptability of their approaches (Moreno-Garcia et al., 2023; Pilz et al., 2024). However, the majority of the studies acknowledged the need for further validation and fine-tuning of the models when applying them to new domains or research questions.
Several limitations and barriers to the implementation of NLP techniques for abstract screening were identified in the included studies. One of the main challenges is the availability of high-quality, labeled training data. The performance of supervised machine learning models heavily relies on the quality and representativeness of the training dataset. Studies with smaller training sets (Bravo et al., 2021) or imbalanced class distributions (Campos et al., 2024) reported lower performance metrics compared to those with larger and more balanced datasets.
Another barrier is the need for domain expertise and technical skills in NLP and machine learning. The development and implementation of NLP models require a deep understanding of the underlying algorithms, feature engineering techniques, and evaluation metrics. Collaboration between subject matter experts and NLP practitioners is crucial for ensuring the validity and relevance of the models in the context of evidence synthesis (O'Mara-Eves et al., 2015). We did not independently test the NLP methods on original or our own data. Hands-on evaluation could provide insight into practical usability and reproducibility for non-specialist users and is recommended for future research.
The computational requirements and infrastructure costs associated with implementing advanced NLP techniques, such as deep learning models, can also pose challenges, particularly for resource-constrained research teams or institutions. While some studies utilized open-source software and libraries (Moreno-Garcia et al., 2023; Pilz et al., 2024), others relied on proprietary tools or commercial platforms (Gartlehner et al., 2019; Hamel et al., 2020), which may limit their accessibility and reproducibility.
User-friendliness and ease of use are important factors for the adoption of NLP tools by systematic reviewers and other end-users. Some studies developed user interfaces or integrated their models into existing software platforms (Manion et al., 2024; Perlman-Arrow et al., 2023) to facilitate their use by non-technical users. However, the majority of the studies focused primarily on the technical aspects of the NLP models and did not provide detailed evaluations of their usability or user acceptance.
These findings have important implications for review teams considering the use of NLP automation tools. Implementation requires careful planning and, ideally, consultation with machine learning experts to select appropriate modeling approaches, preprocess and transform textual data, construct representative training and validation sets, and secure requisite computing infrastructure. When resources permit, starting with a high-performing pretrained model like a fine-tuned BERT variant is a reasonable strategy. However, active learning techniques that iteratively update models as screening progresses may offer a more sample-efficient approach, particularly for novel or niche review topics where seed training data is limited. In all cases, reviewers must remain engaged in the process, providing oversight and high-level direction to ensure the credibility and validity of the machine learning assistance.
A key strength of the evidence base is the generally low risk of bias of the included studies based on the formal ROBIS assessment. The majority of studies satisfied key quality criteria related to the specification of eligibility criteria (Domain 1), study identification and selection processes (Domain 2), and the appropriateness of analytic methods (Domain 4). Some studies lacked clarity in their reporting of data collection and study appraisal procedures (Domain 3), but none were found to be at high risk of bias in any domain.
However, several limitations of the evidence synthesized in this review should be noted. Inconsistencies in datasets, algorithms, implementation choices, and performance metrics across studies precluded quantitative meta-analysis to obtain more precise estimates of expected accuracy and workload savings. The heavy reliance on retrospective evaluations using convenience datasets derived from previously completed reviews means that the ability to generalize results to prospective use cases remains largely untested. Additionally, the brisk pace of progress in machine learning, and especially natural language processing, means that even relatively recent studies can quickly become outdated as new modeling architectures and training paradigms emerge. Regular updates will be necessary to keep pace with this rapidly advancing field. We note that our search was conducted in December 2024, and the period since has seen an acceleration in published research on LLM-based and other novel AI approaches to screening. The findings of this review should therefore be interpreted as reflecting the state of evidence through that date, and readers are encouraged to consult emerging literature for more recent developments. An updated review is planned.
This review benefits from several methodological strengths, including a comprehensive search strategy encompassing multiple bibliographic databases and gray literature sources; adherence to current best practice guidelines for the conduct and reporting of systematic reviews; the use of a formal, multi-domain risk of bias assessment tool; and dual independent study selection, data extraction, and quality appraisal to minimize error and bias.
However, some limitations should be acknowledged. The scope of this review was restricted to applications of NLP techniques for title and abstract screening. Additional research is needed to characterize the utility of these technologies for subsequent labor-intensive steps in the systematic review process, such as full-text screening and data extraction. The rapid emergence of large language models for screening tasks since our search period means relevant LLM-based studies may not have been captured, and future updates should explicitly incorporate this literature.
Our findings highlight several priority areas for future research to advance the development and uptake of NLP tools for systematic review automation. First and foremost, the field would benefit immensely from the creation and sharing of high-quality, standardized benchmark datasets that could serve as a common reference for model development and evaluation. Ideally, such datasets would span a range of research domains and publication types, and would be sufficiently large to support robust training and validation of complex models. Living systematic review datasets that are continuously updated as new evidence emerges could be especially valuable in this regard.
Future studies should prioritize prospective, real-world evaluations of NLP models in the context of newly initiated reviews. Such pragmatic evaluations are essential to establish the external validity and generalizability of tools developed and tested using convenience datasets. In addition to overall accuracy and workload reduction, these studies should explicitly assess risks of bias and the reliability of performance across diverse literature sources, including research from low- and middle-income countries and other underrepresented populations.
As the technical performance of NLP models continues to improve, research efforts should increasingly focus on user-centered design and implementation issues to ensure that these powerful technologies can be readily adopted and used effectively in practice. This includes the development of intuitive, user-friendly interfaces that integrate seamlessly with existing systematic review management platforms and accommodate varying levels of technical expertise. Explainable and interactive AI techniques that facilitate transparency, human interpretability, and appropriate user control and oversight will also be key to fostering trust and promoting responsible use.
Beyond these empirical lines of inquiry, this review underscores the need for deeper interdisciplinary collaboration and exchange between the systematic review and computer science/informatics communities. Forging a shared understanding of the values, goals, and constraints that shape each domain will be critical to ensuring that novel NLP tools are not only technically sophisticated, but also well-aligned with the real-world needs and challenges of evidence synthesis. Building a robust pipeline of researchers with hybrid expertise in both systematic review methodology and machine learning, and establishing open channels for bidirectional knowledge transfer and co-production, will help to bridge current gaps and accelerate progress towards truly transformative innovation in this space.
Authors’ conclusions
The results of this systematic review indicate that NLP technology shows considerable potential for enhancing the efficiency of evidence syntheses, though the wide variability in reported performance (workload reductions 13–96%; recall 14–96%) indicates that benefits are highly context-dependent and not yet consistently achievable across settings. The substantial range in performance metrics underscores that NLP-assisted screening is not yet a reliable ‘plug-and-play’ solution. Performance depends heavily on dataset characteristics, inclusion prevalence, model selection, training data quality, and domain specificity. To date, the application of cutting-edge deep learning models, particularly those that leverage transformer-based language model pretraining, have demonstrated the greatest potential for accurate screening prioritization and workload reduction. However, significant challenges related to accessing high-quality training data, adapting models to new research areas, furnishing necessary computational resources and technical expertise, and ensuring unbiased, trustworthy, and user-centered implementations remain to be addressed.
Realizing the full potential of NLP technologies to revolutionize the speed and scale of evidence syntheses will require active collaboration between systematic reviewers, computer scientists, software engineers, and other key stakeholders to co-develop and rigorously evaluate next-generation tools that are reliable, generalizable, interpretable, and thoughtfully integrated with human expertise and workflows. With sustained research and development along these lines to establish clear best practices and reporting standards, NLP-assisted systematic review pipelines could dramatically improve the timeliness and comprehensiveness of the evidence needed to inform high-stakes health and policy decisions. Ultimately, bridging the power of artificial and human intelligence in this way could have transformative impacts on the real-world usefulness and uptake of research evidence synthesis.
Supplemental Material
Supplemental Material - Techniques, Performance, and Feasibility of Natural Language Processing for Abstract Screening in Evidence Synthesis: A Systematic Review
Supplemental Material for Techniques, Performance, and Feasibility of Natural Language Processing for Abstract Screening in Evidence Synthesis: A Systematic Review by Ravi Shankar, Ziyu Goh, Qian Xu in Campbell Systematic Reviews.
Footnotes
Acknowledgements
We acknowledge the Campbell Coordinating Group for their guidance on systematic review methodology and reporting standards.
Authors Contributions
The review team includes members with complementary expertise in systematic review methodology, natural language processing, and clinical research:Ravi Shankar (RS) - Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Writing – original draft, Writing – review & editing: Content and methodological expertise in systematic reviews and evidence synthesis, with experience in healthcare research and innovation. Ziyu Goh (ZG) - Data curation, Investigation, Validation, Writing – review & editing: Methodological expertise in systematic review methods and medical research, with training in evidence-based medicine. Qian Xu (QX) - Methodology, Supervision, Validation, Writing – review & editing: Clinical and methodological expertise in respiratory and critical care medicine, with experience in evidence synthesis and systematic review methodology. The team composition includes content expertise in healthcare and systematic review methodology (all authors), methodological expertise in systematic reviews (RS, ZG, QX), and information retrieval expertise through consultation with library specialists. Statistical expertise was provided through the narrative synthesis approach appropriate for the heterogeneous nature of the included studies.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
All data supporting the conclusions of this systematic review are included within the manuscript and its tables. The complete dataset of included studies, their characteristics, extracted performance metrics, and risk of bias assessments are presented in Tables 1–3. The PRISMA flow diagram (
) details the study selection process with specific numbers at each stage. All data extraction was conducted using a standardized form based on the methodology described in the Methods section, and no additional datasets were generated or analyzed during this study beyond what is reported in the manuscript.
Plans for Updating This Review
This review will be updated every 3-4 years or sooner if significant new evidence emerges that could materially change the conclusions. RS will be responsible for coordinating updates, with support from the original review team. Given the rapidly evolving nature of artificial intelligence and natural language processing technologies, we will monitor the literature annually to assess whether earlier updates are warranted. Updates will follow the same methodological approach as the original review, with search strategies adapted to include new databases and terminology as the field evolves. The updated review will be submitted to Campbell Systematic Reviews to maintain continuity and ensure accessibility to the evidence synthesis community.
Differences Between Protocol and Review
This review followed the pre-registered protocol (PROSPERO CRD42024615153) with several modifications to enhance comprehensiveness, methodological rigor, and transparency of reporting. The key deviations are described below. First, the search period was extended from August to December 2024 to capture additional recent publications given the rapid development in natural language processing. Despite this extension, we acknowledge that the search is now more than one year old at the time of peer review (April 2026). During this interval, the AI and NLP literature — particularly research involving large language models for screening tasks — has expanded substantially. Ideally, a search update would be conducted; however, changes in the review team’s institutional circumstances, including loss of access to several database subscriptions, preclude a meaningful re-execution of the search at this time. A simple re-run of an amended strategy would also be insufficient given the volume of new literature; a comprehensive update incorporating new search results, screening, data extraction, and synthesis would effectively constitute a new review cycle. We have therefore elected to present the current findings as a synthesis of the evidence through December 2024 and to transparently acknowledge this as a limitation. We note that our protocol commits to updating this review every 3–4 years or sooner if significant new evidence emerges; we believe the threshold for an early update has been met and plan to initiate this process. Readers should interpret our findings in the context of the search date and consult emerging literature for more recent developments. Second, in response to peer review, we acknowledge several limitations in the original search strategy that may have reduced sensitivity. These include the absence of potentially relevant terms such as “active learning,” “citation screening,” “technology-assisted review,” “prioritisation,” “prioritization,” “relevance feedback,” “research synthesis,” “knowledge synthesis,” and “evidence and gap map”; missing plural forms for key concepts (e.g., “evidence syntheses,” “systematic literature reviews,” “meta-analyses,” “scoping reviews,” “rapid reviews”); overly restrictive phrase searching for the screening concept that would have missed variant phrasings (e.g., “title and abstract screening,” “title/abstract screening,” “screening of titles and abstracts”); and limited use of subject headings despite the original text stating otherwise — the PubMed strategy did not incorporate MeSH terms such as “Natural Language Processing” [MeSH] or “Systematic Reviews as Topic” [MeSH], and the use of phrase searching would have disabled automatic term mapping. For the reasons outlined above, it was not feasible to re-run an amended search. To partially mitigate these gaps, we conducted backward and forward citation searching of all included studies and relevant review articles, searched multiple grey literature sources and preprint servers, and consulted content experts, which may have captured some studies missed by the database searches. The sentence in the Methods stating “The search utilized keywords and subject headings” has been corrected to “The search utilized keywords and, where available, subject headings (e.g., MeSH terms in PubMed)” to accurately reflect what was implemented. Third, backward and forward citation searching and expert consultation were conducted as supplementary search methods but were not pre-specified in the PROSPERO protocol. These additions were made to enhance search comprehensiveness in line with best practice guidance for systematic reviews of methods research. Fourth, we employed an adapted version of the ROBIS tool (Whiting et al., 2016) rather than the standard version. The original ROBIS tool was designed to assess risk of bias in systematic reviews of clinical interventions and required adaptation for methodological studies evaluating NLP tools. The four signaling domains were adapted as follows: Domain 1 (Study Eligibility Criteria) assessed whether the datasets used were relevant and clearly defined; Domain 2 (Identification and Selection) assessed whether studies adequately described their data sources and selection processes; Domain 3 (Data Collection and Study Appraisal) assessed the completeness of reporting on feature engineering, model parameters, and validation procedures; and Domain 4 (Analysis and Synthesis) assessed the appropriateness of analytic methods and completeness of results reporting. The adapted tool is provided in Supplemental Table S2. Fifth, additional searches were conducted in conference proceedings and preprint servers (medRxiv, bioRxiv, arXiv) that were not originally specified in the protocol but were deemed important for capturing cutting-edge research in this rapidly evolving field. Sixth, while we initially planned to explore quantitative meta-analysis, the substantial heterogeneity in study designs, datasets, algorithms, and outcome measures made this inappropriate. We instead conducted a comprehensive narrative synthesis following SWiM guidelines (Campbell et al., 2020) as specified in our protocol. Seventh, in response to peer review, the reporting of search methods has been improved. Complete search strategies for all databases, including grey literature sources, the platform and sub-databases searched, dates of each search, and the number of records retrieved, are now provided in Supplemental Appendix S1, following the PRISMA Search Extension (PRISMA-S) reporting recommendations (Rethlefsen et al., 2021). The eligibility criteria section has been expanded to provide a formal definition of NLP for inclusion purposes and to clarify the scope of eligible automation approaches, conference paper eligibility, and the requirement for empirical evaluation. These modifications did not alter the fundamental research questions or inclusion criteria but were implemented to ensure the review’s methodological appropriateness, completeness, and transparency of reporting.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
