Optimized Text Data Processing
Saransh Gupta · 2018 International Conference on Advances in Computing, Communication Control and Networking (ICACCCN) · 2018
As the data generated by the average user continues to rise day by day, the ability to process these large chunks of data seems to be unable to match it. This is mostly caused due to algorithmic inefficiency and time and memory constraints. While numerous algorithms have been there since a long time in the past, very few have ventured texts as a non linear collection of data, an idea that could possibly increase our `large string' processing capabilities by manifold times. If a string is a million characters long, processing it linearly is not nearly the best approach to process it. Algorithms like Rabin Karp and Knuss-Morris-Pratt may also prove inefficient when a large number of operations are needed to be performed on very large texts. In programming languages like Python where strings are immutable in nature, a very large memory is also consumed by such processes. The idea mentioned in the paper helps to deal with such cases in a highly efficient manner in which the efficiency increases along with the size of the data thus making it highly suitable for very large string based data.