An Efficient framework for Metadata Extraction over Scholarly Documents using Ensemble CNN and BiLSTM Technique
P Raghavendra Nayaka, Rajeev Ranjan · 2023
The conventional text documents have made it possible to efficiently retrieve large amounts of text data with the development of various search engines. However, these traditional search approaches frequently have lower accuracy in retrieval, particularly when documents have certain characteristics that call for more in-depth semantic extraction. A search engine for algorithms called Algorithm Seer has recently been developed. The normal search engine collects the deep textual metadata and pseudo-codes from research papers. However, such a system is unable to accommodate user searches that attempt to identify algorithm-specific information, such as the datasets on which algorithms operate their effectiveness, runtime complication, etc. A number of improvements to the previously suggested algorithm search engine are given in this study. We provide various ways to identify automatically and extract pseudo-codes and phrases which transmit metadata utilizing various machine learning methods. Around the 89,000 text lines are used for conducting the experiments; we provided new properties to extract algorithmic pseudo-codes. These characteristics include feature groups with a focus on content, font style, and structure. Our suggested pseudo-code extraction method outperforms current strategies by 28% and obtains a 94.23% Classification Accuracy. Additionally, we suggest a technique for extracting phrases linked to algorithms utilizing deep neural networks, which achieves an 82% of accuracy compared to recent Rule–based provides 23.5% and support vector machine provides 21.5%.