pytorch-widedeep: A flexible package for multimodal deep learning
J. Rodríguez Zaurín, Pavol Mulinka · The Journal of Open Source Software · 2023
In recent years datasets have grown in size and diversity, combining different data types.Multimodal machine learning projects involving tabular data, images and/or text are gaining popularity (e.g.Garg et al. ( 2022)).Traditional approaches involved independent feature generation from every data type and their combination in the later stage before passing them to an algorithm for classification or regression.However, with the advent of "easy-to-use" Deep Learning (DL) frameworks such as Tensorflow (Abadi et al., 2015) or PyTorch (Paszke et al., 2019), and the subsequent advances in the fields of Computer Vision, Natural Language Processing or Deep Learning for Tabular data, it is now possible to use state-of-the-art DL models and combine all datasets early in the process.This has two main advantages: (i) we can partially or entirely skip the feature engineering step, and (ii) the representations of each data type are learned jointly.This means that such representations contain information: (i) related to the target if the problem is supervised; and (ii) how the different data types relate to each other.Furthermore, the flexibility inherent to DL approaches allows the usage of techniques primarily designed only for text and/or images to tabular data, e.g., transfer learning or self-supervised pre-training.With that in mind, we introduce pytorch-widedeep, a flexible package for multimodal deep learning designed to facilitate the combination of tabular data with text and images.