Implementation of a Databaseless Web REST API for the Unstructured Texts of Migne's Patrologia Graeca with Searching capabilities and additional Semantic and Syntactic expandability

Evagelos Varthis, Marios Poulos, Ilias Yarenis, Sozon Papavlasopoulos · 2019

The search through large corpora of unstructured text on the Web Domain is not an easy task and such services are not offered for the common user. One such corpus is the published works of east Christian fathers by Jacques Paul Migne, known as Patrologia Graeca (PG). In this paper, an application of a Databaseless model is presented for extracting information from the unstructured patristic works of PG on the Web Domain. The user queries terms that may exist in PG and retrieves all the paragraphs-fragments that contain these terms. The time for retrieving the information is faster than implementing this querying system in a common Relational Database Management System (RDMS). Our proposed system is portable, secure and can be easily maintained by institutions or organizations. The system auto-transforms the PG corpus into a Representational State Transfer Access Point Interface (REST API) for retrieving and processing the information, using the JavaScript Object Notation (JSON) format. The User Interface (UI) is completely distinguished from the backbone of the system while the system, on the other hand, can be easily extended for more complicated queries as well as to be applied to other corpora. Two kinds of user interfaces are described: The first one is completely static, useful for the average user using the Web browser. The second one illustrates the use of simple Shell Scripting, for searching and extracting statistical, syntactical and semantic information in real time. In both cases we try to strip down the complexity in order to accomplish the corpus transformation and searching, in a simple, secure and manageable way. Difficulties and key problems are also discussed.

Read the paper · More papers on PaperTik