PentaCMD: 299,329 English-to-terminal-command pairs across five tool families, with a leak-free evaluation split
Debnath, Suman · Zenodo (CERN European Organization for Nuclear Research) · 2026
A dataset of 299,329 instruction-command pairs mapping plain-English requests to single terminal commands across five tool families: bash, git, npm, python and PowerShell. Each record is {"instruction", "command", "family"}, where family is supplied as an input hint rather than predicted — identical English maps to different commands across shells ("go to the src folder" is `cd src` in bash and `Set-Location src` in PowerShell). Composition: git 120,000 (synthetic); python 90,000 (synthetic); bash 41,329 (NL2Bash corpus plus synthetic beginner coverage); PowerShell 28,000 (synthetic); npm 20,000 (synthetic). LEAK-FREE SPLIT The dataset ships with a train/validation/test split (269,396 / 14,967 / 14,966) built to eliminate paraphrase leakage, which is the main validity threat for this task. Two different instructions — "install react" and "install the react package" — resolve to the same command. Split those at record level and the model has effectively seen the answer before it is asked the question. Splits are therefore built at group level: records are grouped by target command, groups sharing an identical instruction string are merged, and each group is kept whole within a single split. Stratified per family at 90/5/5. Verified on the released files: no command and no instruction string appears in more than one split. PROVENANCE AND LICENSING The bash subset incorporates the NL2Bash corpus (Lin, Wang, Zettlemoyer and Ernst, LREC 2018; arXiv:1802.08979), which is distributed under the MIT License, Copyright (c) 2020 NL2Bash dataset. That notice is included as LICENSE-NL2BASH.txt. All other families are synthetic, generated from templates for this project. This dataset was built to train PentaCMD-47M, a 47.2M-parameter decoder-only transformer, but is released independently and is usable for any instruction-to-command work.