Analyzing Deep-Learning Kernel Statistics Through Timm
Marika E. Schubert, David Langerman, Calvin B. Gealy, Evan W. Gretok, Alan D. George · 2025
The field of deep-learning models for computervision tasks has evolved rapidly over the past decade, resulting in a wide variety of models with different computational and memory needs. Before committing to a particular model, engineers must understand the general requirements of models in their domain to choose the proper hardware and algorithms. To identify the constraints for a deep-learning model, this paper proposes deconstructing a repository of models in a given domain down to kernels. These kernels then provide insight into general needs for optimization or acceleration in future research. This method is applied to a repository of deep-learning models for image classification. The models chosen are from the PyTorch-based Timm repository, a widely used library containing hundreds of well-maintained reference implementations of imageclassification models. To this end, we built the Timm Crawler (TiC) to aggregate data across this repository. Leveraging this parser, we present a survey of the models contained in this repository relative to kernels used, size of models, and other trends in the field as a whole through the lens of Timm. A key insight from this research is that while the size of the largest models is growing, the median size is not. Within the Timm repository, 92.71 percent of Conv2d layers use$1 \times 1$and$3 \times 3$convolutional kernels. Linear layers are overwhelmingly implemented with biases (91.53 percent) and do not tend to have input/output dimensions that are powers of two (33 percent and 25 percent respectively). These insights into construction indicate that when looking to perform image-classification tasks, there should be a focus on optimizing or designing accelerators for these types of kernels.