Checksums and the internet

David R. Cheriton, Craig Partridge, Jonathan Richard Stone · 2001

This dissertation examines errors above the link layer in data communications networks. Specifically we study the network and transport layers of the Internet protocol suite. The focus of our study is error-detecting codes (EDCs), such as checksums and cyclic redundancy checks (CRCs); how often those EDCs are called upon to detect errors, and how well they perform that task. As a framework for the empirical work in the thesis, we develop a formal description of how to compare and rank different EDCs. We show that currently deployed EDCs (the 16-bit TCP checksum, and the alternate Fletcher checksum) should come close to the ideal of 1 in 216 undetected errors. We present herein a set of simulation studies for existing EDCs which show that in practice these EDCs do not approach the expected ideal results. Our analysis shows that the cause is self-correlation of the input data. In this case, relocating the checksum value to the end of the packet gives a significant improvement. Errors captured from live networks are also examined. We present a new method for capturing both damaged packets from real-life networks and subsequent retransmissions of the damaged packets. This method meets privacy and confidentiality constraints on interception of real-life user data. A significant body of errors was captured at four representative sites. This data is described, summarized, and analyzed. Comparison of the damaged packets and the undamaged retransmissions yields information on the kinds and causes of errors which occur above the link layer. Our data provides strong evidence that errors above the link-layer are due to a variety of causes, which include bad memory in intermediate routers, and bugs in end-host hardware or software which are outside the protection of link-level error checks. For the real-life captured errors, we find CRCs are little or no better than additive checksums of the same size. Our data shows that the datapath between packet buffers and on-NIC CRC hardware is a significant source of errors. We conclude that reliable data communication should use end-to-end error checks. Strongly reliable communication should use an application-level end-to-end check that also covers the packet-buffer-to-NIC datapath.

Read the paper · More papers on PaperTik