Towards the automated extraction and refactoring of NoSQL schemas from application code
Carlos Javier Fernández Candel, Anthony Cleve, Jesús Garćıa Molina · Journal of Systems and Software · 2026
• A code analysis approach to extract NoSQL logical schemas with both reference and aggregation relationships, using a generic metamodel (U-Schema) suitable for multimodel database tools. • The approach enables the automation of database refactorings, including join elimination through field duplication to improve query performance. • The process relies on platform-independent metamodels to represent code, control flow, and data operations, ensuring extensibility and language-agnostic reuse. • The approach incorporates dedicated algorithms for model transformation, from control flow extraction to schema inference and refactoring plan generation. Most NoSQL systems adopt a schema-on-read approach to promote flexibility and agility: the structure of the stored data is not constrained by predefined schemas. However, the absence of explicit schema declarations does not imply the absence of schemas themselves. In practice, schemas are implicit in both the application code and the stored data, and are essential for building tools such as data modelers, query optimizers, data migrators, or for performing database refactorings. As a result, NoSQL schema inference (also known as schema extraction or discovery) has gained attention from the database community, with most approaches focusing on extracting schemas from data. In contrast, the source code analysis remains less explored for this purpose. In this paper, we present a static code analysis strategy to extract logical schemas from NoSQL applications. Our solution is based on a model-driven reverse engineering process composed of a chain of platform-independent model transformations. The extracted schema conforms to the U-Schema unified metamodel, which can represent both NoSQL and relational schemas. To support this process, we define a metamodel capable of representing the core elements of object-oriented languages. Application code is first injected into a code model, from which a control flow model is derived. This, in turn, enables the generation of a model representing both data access operations and the structure of stored data. From these models, the U-Schema logical schema is inferred. Additionally, the extracted information can be used to identify refactoring opportunities. We illustrate this capability through the detection of join-like query patterns and the automated application of field duplication strategies to eliminate expensive joins. All stages of the process are described in detail, and the approach is validated through a round-trip experiment in which a application using a MongoDB store is automatically generated from a predefined schema. The inferred schema is then compared to the original to assess the accuracy of the extraction process.