Statistical methods to improve the accuracy of retrieving similar images in histopathological information database
Yuichi Ishibashi, Tsushima-naka Okayama · 2011
We have developed a histopathological information database system that enables retrieval of similar images. At present, the database stores the images of histopathological specimen including medical certificate text information, that are related to breast disease. The system retrieves cases which resemble the target image and calculates the probabilities for possible diseases connected to the target. Digitization of images is necessary for image retrieval, therefore large specimen images are divided into many small images and Wavelet transformation is performed upon each small image. The small images that characterize the diseases are taken as training data, and the divided small images are identified by pattern recognition by the Neural Network. The result of this identification is used as a feature vector for a specimen image. Similar images are retrieved by comparing the feature vectors of the targeted image and the specimen images in the database. A medical certificate describes the histopathological diagnostic process which is completed by pathological doctors. This diagnostic information is digitally transformed by the extraction of keywords and creates a dictionary by applying a text mining technique. The types of keywords that are involved in differential diagnosis of diseases are analyzed by canonical discriminant analysis, using the disease as a group and the keyword as a variable. This result enables the calculation of probabilities for possible diseases by Bayes' theorem including the keywords. Furthermore, the probabilities for possible diseases of the targeted image can be calculated by combining training data and keywords that represent training data. In order to improve the accuracy of similar image retrieval, it is necessary to select appropriate training data. Characteristic images are selected as training data referencing the keywords that associate with the diseases mentioned in the above analysis. Interestingly, the principal component analysis detected a unique group of training data that resembled the other types or a distant data of the same category. Consolidation of training data and the addition of the deficient data with specific information, assisted the decrease in the misclassification of pattern recognition. 1. The method for the calculation of image feature vectors and image retrieval