The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and Generation

Dung Nguyen, Le Nam, Anh T. V. Dau, Anh Nguyen, Khanh Nghiem, Jin Guo, Nghi Bui · 2023

We present The Vault, a dataset of high-quality code-text pairs in multiple programming languages for training large language models to understand and generate code.We present methods for thoroughly extracting samples that use both rule-based and deep learningbased methods to ensure that they contain highquality pairs of code and text, resulting in a dataset of 43 million high-quality code-text pairs.Our extensive evaluations on common coding tasks including code generation, code search and code summarization show that when fine-tuning Code Large Language Models on The Vault, such models outperform the same models trained on other datasets such as Code-SearchNet.We also provide detailed analyses of our datasets to assess the effects of various programming languages and docstrings on the performance of such models.

Read the paper · More papers on PaperTik