A Curated Set of Labeled Code Tutorial Images for Deep Learning

Adrienne Bergh, Paul Harnack, Abigail Atchison, Jordan Ott, Elia Eiroa-Lledo, Erik J. Linstead · 2020

As more work is conducted in mining software repositories at a large scale, aggregating data from multiple modalities beyond textual representations will prove beneficial. Online video content related to software engineering offers another, thus far largely untapped, venue for source code. In this paper, we present a novel dataset comprised of 111,229 frames extracted from a sample of 100 technology-related YouTube tutorials, taught in the Java and Python programming languages. These videos span a diverse range of Integrated Development Environments (IDEs) and modes of information presentation, lending heterogeneity of image features to our dataset. Hand-labeled for the presence, or lack thereof, of typeset, partially visible typeset, or handwritten code, the dataset provides a uniquely consistent and accurate data source for supervised machine learning ventures in extracting and processing code from images.

Read the paper · More papers on PaperTik