Maintenance of File System And Improving Efficiency Of Hadoop By Using Clustering
Samandeep Kaur, Kewal Krishan · 2013
In the existing work, Hadoop distributed file system is defined as the large cluster set in which number of servers is collected for direct storage of data. The file system is defined as architecture for large dataset. In this present work we are basically maintaining the file system in form of clusters along with respective mount table definition. The mount table is attached with each cluster and the optimization of user query is been performed based on same mount table. The work is extended in two main phases, in very first phase the Hadoop architecture is defined in clusters. The cluster formation is an intelligent formation on keyword based feature analysis on files. The related files are kept in one cluster. Along with this, the mount table is defined that stores the keyword information as well as other metadata related to each file over the system. Once the architecture is defined, in the second phase the user query is filtered and the keyword is extracted from it. Based on the keyword analysis, the related cluster is selected and the query is performed on the selected cluster. The proposed work will improve the efficiency of the system.