Reference-Corpus Formation for Estimating the Closeness of Topical Texts to the Semantic Standard
D. V. Mikhaylov, G. M. Emelyanov · Pattern Recognition and Image Analysis · 2022
Abstract Estimation of the closeness of a topical text to the most rational (i.e., standard) form of expression of the sense corresponding to it in the given natural language requires the existence of some representative reference collection (or corpus) of texts that reflect the most significant concepts of a given topical area and relationships between them. In the current study, the problem of forming such collection is solved by using the occurrence of words from abstracts of articles, selected by an expert, in documents under estimation for inclusion into the reference corpus. The estimation is based on the comparison of values for the 5-th percentile of the empirical distribution corresponding to an array of fractions for nonzero values of the term frequency (TF) for separate phrases within each abstract relative to the analyzed document. Herewith, the most significant documents for the reference corpus will be those with the maximal value of the mentioned percentile. Experiments show that the proposed solution gives at least a fivefold reduction in the number of documents of those that are minimally relevant to a given topical area when implementing selection for the reference corpus.