Evaluating Linear Shallow Autoencoders on Large Scale Datasets

Vojtěch Vančura, Petr Kasalický, Rodrigo Augusto da Silva Alves, Pavel Kordík · ACM Transactions on Recommender Systems · 2025

Linear shallow autoencoders have gained popularity in collaborative filtering benchmarks for recommender systems due to their simplicity and strong performance. However, models like EASE struggle to scale to datasets with a large number of items. This paper addresses this limitation by evaluating scalable variants of shallow linear autoencoders, namely, the recently proposed ELSA and SANSA, on large-scale datasets representative of real-world applications. We further propose improvements to enhance their scalability on large, sparse datasets. Our evaluation goes beyond standard offline and online performance metrics, incorporating key industrial considerations such as training time and memory usage. By analyzing the effects of data sparsity and catalog size, we provide practical guidance to select the most suitable model for a given dataset. We release the source code, training logs, and detailed instructions to reproduce our experiments at https://github.com/recombee/SparseELSA.

Read the paper · More papers on PaperTik