Large Language Models in Software Project Management: A Two-phase Systematic Literature Review

Muhammad Saad Amin, Emilia Mendes, Adam Alami, Ricardo Britto · Journal of Systems and Software · 2026

Background Large Language Models (LLMs) are increasingly integrated into Software Engineering (SE) processes and practices. However, their application within Software Project Management (SPM) remains fragmented and insufficiently synthesized. This absence of systematic evidence prevents both researchers and practitioners from understanding where LLMs can meaningfully contribute to project management processes in relation to established standards such as the Project Management Body of Knowledge (PMBOK). Objectives This study aims to systematically synthesize existing evidence on the application and integration of LLMs in SPM by means of a two-phase method where Phase I systematically searches for existing secondary studies in the topic of interest, and Phase II (whenever applicable) carries out an SLR in the topic of interest. Method We conducted a two-phase method: Phase I is a tertiary study that used 14 research questions to assess the existence, coverage, and quality of prior secondary studies in LLMs applied to SPM. As no suitable SLR was identified from Phase I, we also conducted Phase II, which employed the SEGRESS guidelines to analyze primary studies that empirically evaluated the use of LLMs in SPM via five research questions. Phase II also assessed coverage across PMBOK processes and synthesized reported mechanisms, outcomes, research designs and evaluation approaches. Results The SLR (Phase II) shows that studies employ encoder-only and decoder-only LLM architectures in equal proportion (17 studies each). Their applications focus on a narrow set of SPM tasks, including effort estimation, project planning, and requirements-related activities. Coverage across PMBOK processes is uneven, with limited attention to areas such as procurement, stakeholder, quality, and cost management. Most studies rely on proof-of-concept evaluations, with scarce industrial validation and heterogeneous evaluation metrics. Conclusion The results indicate that current research on LLMs in SPM provides limited guidance for systematic adoption. One of the main drawbacks for such lack of guidance is the weak empirical validation in many of the included primary studies, limited PMBOK coverage, and lack of context-aware integration of LLMs in SPM. 1

Read the paper · More papers on PaperTik