Frequency Value Grammar and Information Theory
Asa M. Stepak · Journal of Applied Sciences · 2005
This paper convincingly shows that previous efforts to calculate the Entropy of written English are based upon inadequate n-gram models that will result in an overestimation of the Entropy.Frequency Value Grammar, FVG, is not based upon corpus based word frequencies such as in Probability Syntax, PS.The characteristic 'frequency values' in FVG are based upon 'information source' probabilities which are implied by syntactic category and constituent corpus frequency parameters.The theoretical framework of FVG has broad applications beyond that of formal syntax and NLP.In this paper, I demonstrate how FVG can be used as a model for improving the upper bound Entropy calculation of written English.Generally speaking, when a function word precedes an open class word within a phrasal unit, the backward bi-gram analysis will be homomorphic with the 'information source' probabilities and will result in word frequency values more representative of cognitive object co-occurrences in the 'information source'.