The Impact of Incorrect Training Sets and Rolling Collections on Technology-Assisted Review.
Jan Scholtes, T.H.W. van Cann · 2013
Document classification and machine learning technology in electronic discovery (eDiscovery) are gaining attention under new names such as technology-assisted review (TAR), machine-assisted review (MAR), computer-assisted review (CAR) and predictive coding [2]. Several judicial rulings have addressed typical legal concerns in relation to the quality of machine learning. In this paper, we address two of such concerns and investigate their relation with machine learning quality in more detail. The topics of interest of this paper are: (i) the impact of the quality of training documents on the overall classification results, which can be measured by investigating the impact of training supervised classifiers deliberately with wrong training samples and (ii) using machine learning in so-called rolling collections, which can be measured by investigating the quality of classification of new and unknown documents, which have not been used to extract or select machine-learning features, with existing classifiers. A machine learning pipeline with the most-common used document feature-extraction techniques known as a bag-of-word (BoW), and Term Frequency- Inverse Document Frequency (TF-IDF) is used [9]. For feature-selection, vector logarithmic normalization and cut-off of non-used or non-relevant dimensions has been selected. The resulting data was used to train binary