A Big Data architecture for knowledge discovery in PubMed articles

Francesco Gargiulo, Stefano Silvestri, Mario J. Ciampi · 2017

The need of smart information retrieval systems is in contrast with the difficulties to deal with huge amount of data. In this paper we present a Big Data Analytics architecture used to implement a semantic similarity search tool for natural language texts in biomedical domain. The implemented methodology is based on Word Embeddings (WEs) models obtained using the word2vec algorithm. The system has been assessed with documents extracted from the whole PubMed library. It will be also presented a user friendly web front-end able to assess the methodology on a real context.

Read the paper · More papers on PaperTik