On Statistical Sequencing of Document Collections

Ramya Thinniyam · TSpace (University of Toronto) · 2014

This thesis was primarily motivated by the Documents of Early England Data Set (DEEDS), a collection of undated medieval property exchange charters written in Latin. The main goalis to devise automated statistical methods to estimate the temporal order in which a corpus of documents was written using only the words that appear in the corpus. Our interest lies in sequencing the documents, not dating them. The premise is that documents written closer together in time will naturally be more similar in content, and thus we propose the following two-step approach to sequencing: (1) obtain distance measures between pairs of documents (using only their word characteristics); (2) estimate the optimal ordering of the documents based on these distances.We describe various types of distances that can be computed between pairs of documents. We then present three methods for sequencing a set of documents based on their pairwise distances. The first method sequences elements using a regression model on their pairwise distances and optimizes the corresponding Error Sum Of Squares (SSE) to estimate the ordering. The second method called the "Closest Peers" method minimizes the average distance of each document to its closest matching documents. The third is an MCMC approach that not only allows for sequencing, but parameter estimation and natural inferences about orderings on subsets of the documents. The performance of the sequencing methods are evaluated and compared via simulated data.The methods that we describe in this thesis are not only applicable to the DEEDS corpus, but also for other collections such as: drafts/versions of works written by the same author, different documents written by the same author, different documents written on the same topic, etc. Application of the sequencing methods are carried out on real document collections such as the DEEDS corpus and on various sets of drafts. As an addendum, we propose a distance-based method for classifying documents using a training set and apply it to the DEEDS corpus.

Read the paper · More papers on PaperTik