A Distributed Data Allocation Algorithm for Biological Databases
Gustavo Tonini, Frank Siqueira · 2013
Storage and processing of large data sets on distributed platforms allows parallel query execution and is capable of improving scalability. However, defining a distributed allocation schema is a complex task that has been based mostly on ad-hoc, trial-and-error strategies. This paper describes an algorithm for creating a distributed allocation schema aimed at improving query performance. The algorithm is based on data patterns and query history analysis and can be applied to any existing centralized database. The proposed algorithm was evaluated using a large biological database as case study, achieving very promising results.