Computational Optimizations in LLMs

Aishwarya Asesh, Meenal Dugar · 2023

Over recent years, the proliferation of microblogging and textual messaging platforms has led to an exponential surge in textual data, necessitating advanced, automated techniques for sentiment elucidation. Contemporary methodologies, despite their efficacy, often demand substantial computational expenditure and manifest pronounced overfitting in scenarios involving non-standard dataset distributions. Parameter-Efficient Transfer Learning (PETL) has emerged as a promising strategy to mitigate the exorbitant computational overheads associated with large-scale model training, albeit not without its computational burden. This research, an extension of extant literature, introduces a pioneering MINIature Orthogonal Network (MINION) approach. It synergistically couples a diminutive sequence-to-sequence recurrent neural network (RNN) with a static transformer model, eschewing the need for backpropagation through the voluminous transformer. Experimental results affirm that MINION ensures a notable reduction in computational requisites in juxtaposition with antecedent implementations and comprehensive model fine-tuning, simultaneously curtailing model over-assurance while preserving commendable accuracy metrics in sentiment categorization tasks.

Read the paper · More papers on PaperTik