Describing and Accessing Software Engineering Research Knowledge

Angelika Kaplan · Repository KITopen (Karlsruhe Institute of Technology) · 2026

Software Engineering (SE) research drives progress and innovation in science, practice, and society, forming a cornerstone of the digital world.Simultaneously, the digital transformation of the scientific ecosystem is reshaping how research knowledge is acquired, shared, and reused, promoting the principles of Open Science and FAIR (Findable, Accessible, Interoperable, and Reusable).Despite these advancements, scientific papers, typically disseminated as PDF documents, remain the primary medium for communicating research findings.These documents are inherently monolithic and lack machine-readable semantics, which hinders automated processing and effective research knowledge retrieval.In case of SE research, such publications are often further accompanied by supplementary materials, including data and software, which provide the empirical foundation and evidence for validating research claims.These materials are often (if available) linked or stored in a heterogeneous manner, further complicating research data management and the realization of the FAIR principles.This doctoral thesis addresses these weaknesses by developing an approach that semantically describes SE research artifacts and enables access to research knowledge.The approach's architecture, using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG), comprises two core modules: (1) a knowledge description module that supports structured schema representation through semi-automated population of knowledge graphs, and (2) a knowledge access module that enables schema-agnostic retrieval.Central to our approach is a conceptual classification and data schema designed to describe research findings in scientific publications by their validity and supporting evidence.We instantiated this schema in the research field of software architecture and evaluate its feasibility through a literature study of top-tier publications.Furthermore, we propose a general research process based on the GQM approach (Goal, Question, Metric) that integrates descriptive and normative aspects of research artifacts and their knowledge statements, and, thus, supports schema evolution and long-term sustainability.To operationalize our conceptual approaches, we use the FAIR-compliant Open Research Knowledge Graph (ORKG), deriving ORKG templates as blueprints for schema specification and for graph population.We develop a flexible classification framework that supports transfer learning and prompting to train classifiers for semi-automated knowledge graph population based on large language models.The framework is also able to address challenges such as dataset imbalance via data augmentation and text perturbation techniques.The feasibility and maturity of our framework are validated using the labeled dataset of the aforementioned literature study as gold standard dataset for benchmarking.This dataset also serves as the basis to populate the knowledge graph.Besides tracking and assessing conventional performance metrics for the automated classification task, we defined combined sustainability metrics within our framework to enable trade-off decisions between performance and carbon emissions.i

Read the paper · More papers on PaperTik