Vehicle-Scale LLMs: Integrating low-rank residuals and 4-bit quantization for in-vehicle AI

Anubha Parashar, Apoorva Parashar, Apoorva Parashar, Apoorva Parashar, Bhavya Joshi, Kavita Jhajharia, Aditya Sinha · Array · 2026

The proliferation of large language models (LLMs) has enabled transformative advances in in-vehicle conversational artificial intelligence (AI), traffic forecasting, and natural language navigation. However, deploying LLMs on automotive edge devices faces stringent memory and compute constraints, as modern vehicle processors typically provide less than 16 GB of dedicated memory for AI workloads. We present Dynamic Adapter Compensation + 4-Bit Quantization for Intelligent Transportation Systems (DAC+Q4-ITS), a fully training-free compression framework that fuses zero-initialized low-rank adapter compensation with on-the-fly 4-bit dynamic quantization. Our approach begins by collecting a compact in-vehicle corpus comprised of diverse driving-related utterances, traffic reports, and route-planning dialogues. We compute layerwise activation eigenspaces via singular value decomposition (SVD) from this calibration data and project the quantization-induced weight perturbations onto the top eight eigenvectors per transformer layer. This process yields rank-8 adapter compensation modules that are injected directly into each layer without any gradient updates, thereby recovering accuracy lost to aggressive 4-bit quantization. During inference, all linear projections are executed in 4-bit integer format—achieving a fourfold reduction in weight storage—while the corrective adapter paths remain in 32-bit floating-point (float-32). We demonstrate that a 1 billion-parameter (1 B-parameter) LLaMA-derived base model can operate on commodity automotive hardware (e.g., NVIDIA Graphics Processing Unit (GPU) Jetson Xavier NX with 16 GB Low-Power Double Data Rate 4X (LPDDR4x) memory), incurring only a 1%–3% latency overhead relative to an 8-bit baseline. On downstream tasks—natural language navigation (NLN), conversational route planning (CRP), and traffic forecasting (TF)—it achieves over 98 % of full-precision performance. By enabling low-latency, privacy-preserving, on-device LLM inference under strict memory budgets, DAC+Q4-ITS establishes a practical pathway for embedding advanced conversational and predictive AI directly in intelligent transportation systems, eliminating reliance on cloud connectivity and enhancing robustness, reliability, and user privacy.

Read the paper · More papers on PaperTik