Mathematical Models in Information Science

Paul B. Kantor · Bulletin of the American Society for Information Science and Technology · 2002

It has been said that mathematicians are basically puzzle solvers, and that they don't so much care what the puzzles are about. That may be a very good description of my own journey through several parts of information science. When I first encountered this field, in the early 1970s, it was through the application of mathematical models that come from operations research (OR). In work with Tefko Saracevic, now at Rutgers, we developed a kind of branching analysis, which from a mathematician's point of view is a Markov model, to understand better the factors determining the availability of books in a library. At that time, there was already a significant literature about the application of operations research to libraries, stimulated both by the work of the famous founder of OR, Philip Morse, together with Ching-Chih Chen, and Donald Swanson and Abe Bookstein at the University of Chicago. There had also been some earlier work, which bordered on economics, by Fussler and Simon. As I looked at that work, my first thought was that there was nothing left to be done, as the work included fairly sophisticated models for queuing, which is the mathematical phenomenon that seems most to govern availability. However, as I came to know library professionals, I learned that the specific mathematical parameters (those pesky Greek letters) in the formulae were generally unknown. Not only were they unknown, but librarians – not being mathematicians – had no idea how to determine them. Thus, I began a 10- or 12-year detour into the nuts and bolts of performance evaluation, concentrating primarily on developing simple tools that would enable working librarians to estimate the important parameters of an OR model. This work was begun before the introduction of spreadsheets and became substantially more feasible with the invention of VisiCalc. With my collaborator, Jung Jin Lee, now professor of statistics at Soong Sil University in Seoul, Korea, we developed a succession of spreadsheets, eventually developing very complex models using the macro language of the Lotus Symphony spreadsheet. These models, marketed under the name FUNKIEST, provided some moderate stream of income in the late 80s and early 90s. Today of course, every library has its own spreadsheet expert and bends and warps the numbers to suit its own ends. In a parallel development, the availability measures (although not some of the more sophisticated measures of flow and delay) were incorporated into the public library evaluation process, while a richer array was developed for and presented to the academic library community by the Association of Research Libraries. While mathematical interests had been rather peripheral to the central goals of librarianship, I found that they were somewhat closer to the center in modeling human behavior. Although this is a field that seems perpetually in its infancy, we were able to develop (in the 1980s) a kind of Bayesian model for the behavior of a person using any kind of information resource. The central idea is that the person begins with some degree of confidence that the system will provide a useful answer, and this degree of confidence is bolstered, or eroded, by the actual experience of using the system. When the confidence is eroded to the point where the expected value of making one more try at the system has become negative, the user will quit. This model has been adopted by some researchers in industrial settings, but is not yet widely used. More recently, working with Wonsik "Jeff" Shim, now at Florida State University, we explored the application to library evaluation of complex economic models called "Data Envelopment Analysis" (DEA). While the details are quite complex, the essential idea is that DEA captures in mathematics the important fact that libraries are different from each other. Thus no one performance measure can be applied to all of them. But DEA says to a particular library, A, "Devise the measure that makes you look as efficient as is reasonably possible." It then uses that optimizing measure to compare library A with all the other libraries in a large set. If some other library B comes out better, then we suggest that the director and managers of library A make it a point to visit library B and see how they do it. The area of information science where a mathematician, such as myself, finds greatest opportunity for expression is information retrieval (IR). While information retrieval is accomplished by evolving systems, each of which includes a host of algorithms and special support features, there is some reason to believe that it is at its heart mathematically analyzable. The best-known example of such an analysis is the Robertson and Sparck-Jones model, which is a Bayesian model based on the independence of various features that a document might have. When you realize that typically the features that a document has are the presence or absence of particular terms (or in more complicated examples the number of times that particular terms occur) it's pretty hard to believe that they are actually independent. In a series of papers with Jung Jin Lee, we explored more sophisticated models, which in effect assert that the distribution of terms over documents is as independent as it can be, given what we actually know about the distributions of terms. Mathematically this is expressed by adapting the so-called "Maximum Entropy Principle" (MEP) to describe distributions of terms over documents. William Cooper and P. Huizinga had proposed the relevance of the MEP for IR somewhat earlier. What we found, as we looked into it, is that while you can calculate clearly the predicted distribution of relevant documents over various subsets of a collection, based on occurrence of terms, the larger the collection is the worse the predictions of this maximum independence or Maximum Entropy model become. In other words: It didn't work out too well. I remain convinced, however, that statistics and techniques developed in fields ranging from signal analysis to pattern recognition, have a lot to teach us about how to do information retrieval well. As in any fiel d where mathematics is being applied to the problems of the real world, the difficulty is to translate that real world into the correct mathematical expressions. One direction that looks very promising is to work on more sophisticated Bayesian models, such as are being developed by David Madigan and David Lewis in a recently initiated project. Another place to look, somewhat less ambitious, but nonetheless of considerable practical interest, is in the area of data fusion. Working with my former student, and now collaborator, Kwong Bor Ng of City University of New York, Queens, I've looked for a number of years at the question of whether data fusion (which is known to be important in signal detection) can be applied sy stematically to improve the performance of information retrieval systems. The central idea of data fusion is that when several different systems each make estimates of the relevance of a document, knowing the whole sets of estimates provides more information, and should, in principle, enable us to make a more accurate estimate of relevance. Croft and Belkin have explored some kinds of data fusion under the Rubric "Combination of Evidence." Many researchers have explored these techniques typically using linear models based on the scores provided by several systems. With KB Ng, we have looked at the potential upper limit of performance, using non-parametric approaches to decide when fusion would be effective. It has been a great pleasure to see that the evolution of information science has been such as to bring a mathematically oriented person, like myself, ever closer to the center of interest. The years ahead look exciting and I am sure that there is much to be learned. Paul B. Kantor is professor, SCILS, at Rutgers University. He can be reached at 4 Huntington Rd., New Brunswick, NJ 08903; 732-932-1359; [email protected]

Read the paper · More papers on PaperTik