Proximity learning for non-standard big data.
Frank-Michael Schleif · 2014
Abstract. Huge and heterogeneous data sets, e.g. in the life science domain, are challenging for most data analysis algorithms. State of the art approaches do often not scale to larger problems or are inaccessible due to the variety of the data formats. A flexible and effective method to analyze a large variety of data formats is given by proximity learning methods, currently limited to medium size, metric data. Here we discuss novel strategies to open relational methods for non-standard data at large scale, applied to a very large protein sequence database. 1