Nesterov momentum and gradient normalization to improve t-SNE convergence and neighborhood preservation, without early exaggeration

Pierre Lambert, Lee John, Edouard Couplet, Cyril de Bodt · 2023

Student t-distributed stochastic neighbor embedding (t-SNE) finds low-dimensional data representations allowing visual exploration of data sets.t-SNE minimises a cost function with a custom two-phase gradient descent.The first phase is called early exaggeration and involves a hyper-parameter whose value can be tricky and time-consuming to set.This paper proposes another way to optimise the cost function without early exaggeration.Empirical evaluation shows that the proposed method of optimization converges faster and yields competitive results in terms of neighborhood preservation.

Read the paper · More papers on PaperTik