Packet Loss and Duplication Handling in Stream Processing Environment
István Finta, Gergely Elias, Janos Illes · 2018
In the course of transmission through networks a particular packet, like a Storm tuple or a performance/fault management (PM/FM) report in XML format of an Operation Support SystemOSS) application, data loss/out of order arrival/duplication phenomena may cause the packet not to arrive at the destination, to arrive exactly once or to arrive in several copies. Vast literature explores typical distributions for packet loss/duplication. At each layer of the network the entities processing the packets need to cope with packet loss/duplication in a certain way. In some cases loss/duplication does not matter too much. In other cases losses/duplications need to be very strictly handled. At the application level such strict handling is a must in case of an OSS application: missing or duplicated PM data may cause a wrong perception about the real performance of the network. In such cases for every received packet the receiving end needs to check whether the packet has already been received or not and then to insert it exactly once. The efficiency of the checking depends on the data structure being used for storing the packets or their keys. In this contribution we construct a model about typical statistical distribution of packet loss and duplication in stream processing environment. Then through the parameters of the model we assess the performance of an interval-based, lightweight, high speed filtering tree against traditional binary search trees.