Multisubject Analysis and Classification of Books and Book Collections, Based on a Subject Term Vocabulary and the Latent Dirichlet Allocation

Nikolaos Makris, Nikolaos Mitrou · IEEE Access · 2023

In this paper, a new method for automatically analyzing and classifying books and book collections according to the subjects they cover is presented. It is based on a combination of the LDA method for discovering latent topics in the collection, on the one hand, and the description of subjects by means of asubject term vocabulary, on the other. Books, topics and subjects, all are modelled asbag-of-words, with specific distributions over the underlying word vocabulary. TheTable of Contents (ToC)was used to describe the books, instead of their entire body, whilesubject(orstandard)documentsare produced by a subject term hierarchy of the respective disciplines. Frequency-of-terms in the documents and word-generative probabilistic models (as the ones postulated by LDA) were integrated into a consistent statistical framework. Using Bayesian statistics and simple marginalization equations we were able to transform the expressions of the books fromdistributions over unlabeled topics(derived by the LDA) todistributions over labeled subjectsrepresenting the respective disciplines (Physical sciences, Health sciences, Mathematics, etc). More specifically, the necessary theoretical basis is firstly established, with eachsubjectformally defined by the respectivebranch of a subject term hierarchy(much like a ToC) or the respectivebag of words(single words and biwords) produced by flattening the hierarchy branch; flattening is realized by taking all the terms of the nodes and leaves of the branch with repetitions allowed. Being confined within a closed set of subjects, we are able to invert thefrequency-of-termsin each subject [also interpreted as the probability of generating a term (wn) when sampling the subject (si) and denoted by Pr{wn|si})] and express each term as aweighted mixture(orprobability distribution)of subjects, denoted by Pr{si|wn}. This is the key idea of the proposed method. Then, any document (dm) can be expressed as a weighted mixture of subjects (or the respective distribution, denoted by Pr{si|dm}) by simply summing up the distributions of the individual terms contained in the document. This is made possible by virtue of some simple formulas that have been formally proven for the union of documents (Pr{si|(d1∪d2)}) and for the union of subjects (Pr{(si∪sj)|d}). Since not all vocabulary terms are found in a particular set of books, nor, conversely, all corpus words are included in the subject vocabulary either, two important measures come to the foreground and are calculated with the proposed formulation: thecoverage of a book or a corpus by the subject term vocabularyand, conversely, thevocabulary coverage by a set of books. These measures are useful for updating/enriching the subject term vocabulary, whenever it happens that documents with new subjects are included in the corpus under analysis. Following the theoretical formulation, the derived results are combined with the LDA in order to further facilitate our multisubject analysis task: using the subject term vocabulary, LDA is applied on the corpus under study and results in expressing each book (bm), as a probability distribution over hidden topics (denoted by Pr{tk|bm}). In the same framework, each topic (tk) is expressed as a probability distribution over words (Pr{wn|tk}). Having estimated each word’s probability distribution over subjects (Pr{si|wn}), we can express each discovered topic as a weighted mixture of subjects [Pr{si|tk} = ΣnPr{wn|tk} Pr{si|wn}] and, by using that, we express each book in the same manner [Pr{si|bm} = ΣkPr{tk|bm} Pr{si|tk}]. This is a very clear and formal way towards obtaining the desired result. The proposed methodology was applied to a Springer’s e-book collection with more than 50,000 books, while a subject term hierarchy developed by KALLIPOS, a project creating open-access e-books, was used for the proof of concept. A number of experiments were conducted to showcase the validity and usefulness of the proposed approach.

Read the paper · More papers on PaperTik