Reduction Of Dimension In Linear Binary Classification
Osipyants Pavel Maksimovich · Zenodo (CERN European Organization for Nuclear Research) · 2014
Consider the linear problem of binary classification (if the problem is linearly inseparable, it can be led to that by using a symmetric integral L-2 kernel). In solving this problem the classified elements are represented as elements of a vector space of dimension n. In practice n can be extremely large, for example for the task of classification of genes, it can reach tens of thousands. Large dimension implies high computation time and potentially high calculation error. Moreover, the use of large dimension can be expensive (for experimentation). So the question is: how to reduce n by discarding insignificant element components so that the elements were separated "no worse" in new space (were linearly separable) or "not much worse." In this article I want to start with a brief overview of the technique from the article Isabel Guyon, Jason Weston, Stephen Barnhill, Vladimir Vapkin "Gene Selection for Cancer Classification using" and then offer my own method.