Information Retrieval from Electronic Health Records

Meshal Al‐Qahtani, Stamos Katsigiannis, Naeem Ramzan · 2020

Advances in computing encouraged the adoption of computer systems in numerous applications. In the health domain, the adoption of computer systems enables the introduction of better services, the provision of reliable services, and the reduction of human errors. Generally, data in computer systems are stored in coded format. However, in health databases some data cannot be coded, such as doctors, comments; hence, they are stored in the form of free text. Available literature has demonstrated that such free text contains invaluable information. However, extracting information from the free text portion of health databases is a challenging task due to the complexity of the stored data. Latent semantic indexing (LSI) is an information retrieval (IR) technique that has proven its effectiveness in extracting information from health databases, as it is able to identify the semantics of the terms within and across the documents within the database. However, LSI has a major limitation, which is its inefficiency when extracting information from large-scale document collections. In this chapter, two enhancements of the LSI method are proposed and evaluated in order to overcome this limitation. The proposed distributed LSI and parallel LSI methods were applied to an artificial electronic health records database (EMRbots) and were evaluated in terms of time complexity, recall, and precision.

Read the paper · More papers on PaperTik