Predictive use cases of CNN based multi label classification for programming languages

Satyarth Upadhyaya, Anish Parajuli, Subarna Shakya · 2019

Multi-label classification refers to classifying data into two or more, usually independent, set of output labels. This approach is suitable for deep learning applications in multi-faceted subjects like software development, where it is desirable to yield multiple outcomes. This paper proposes a CNN based deep learning model on public datasets of programming language platforms like GitHub and Stack Overflow to infer intelligence to aid decision making process regarding the choice of programming languages for a given software development requirement. For this research, we've developed a training model with pre-trained vector embedding layer and multi-channel one dimensional CNN layers, followed by Multi Layer Perceptron layer to provide multi label outputs. We have managed to achieve 92%, 98% accuracy and 22%, 4% loss with our two experimental setups for Github and Stack Overflow respectively. The model performed well when tested on software development requirements. Stack Overflow dataset was observed to be noticeably better performing than the Github dataset for actual software development use cases. The implications of these models were also found to be good for trend prediction and source code use cases.

Read the paper · More papers on PaperTik