Comparison of Texts Streams in the Presence of Mild Adversaries
Michael Malkin, Ramarathnam Venkatesan · 2005
Text sifting is a method of quickly and securely iden-tifying documents for database searching, copy de-tection, duplicate email detection and plagiarism de-tection. A small amount of text is extracted from a document using hash functions and is used as the document’s fingerprint. We build upon previous work by Broder et al. [4,5] and Heintze [8], specifically ad-dressing a certain set of attacks that we discovered to be very powerful against previous systems. We achieve robustness against these attacks with a new selection process. We also give theoretical and ex-perimental results for these and other attacks on text sifting functions.