Multi-Label Tabular Synthetic Data Generation for Bundle Recommendation Problem

Aakash Swami, V Tirumala · 2023

Recommender systems are a great academic and industrial success in e-commerce and other areas. However, access to real-world historical interaction data for evaluating the recommender systems remains a challenge due to privacy, cost, and multiple issues. The state of the art techniques for generating synthetic data that is closest in relation to real world data, needs atleast a sample of real world data, which is also a problem especially when it comes to bundled data. We propose a novel way of generating multi-label synthetic data specifically for Bundle Recommender system problems, that does not need a sample of real world data but still is closest in relation to real world data. The proposed approach is hybrid as it uses a combination of process based and data based synthetic data generation methods. We demonstrate the feasibility of this approach by generating bundled data for Cosmetics product purchases. Also, synthetic data generated by the state-of-the-art method and our hybrid approach were compared, and it was found that the two sets of data are comparable in terms of statistical characteristics. The generated synthetic data can be used for evaluating bundle recommender systems for the industry.

Read the paper · More papers on PaperTik