Reasoning Beyond Length Limits: Improving Accuracy in Long-Context Question Answering With Small-Scale Language Models

Minyoung Kyoung, Joon-Ho Lim, Youngsoo Kim · IEEE Access · 2025

Long-context question answering (QA) remains a significant challenge, particularly when using small-scale language models (SLLMs) with limited computational capacity. Despite their efficiency, SLLMs often struggle to capture complex reasoning patterns and synthesize information from lengthy documents due to constraints in context size and inference depth. Traditional retrieval-augmented generation (RAG) approaches offer partial relief but typically fall short when precise reasoning across multiple passages is required. In this study, we present a novel, lightweight framework designed to improve long-context QA performance in SLLMs by combining two key strategies: (1) instruction-tuned embedding-based retrieval for extracting semantically aligned context, and (2) a question rephrasing mechanism that decomposes complex queries into stepwise subquestions. This dual strategy enables structured reasoning without the need for additional training or model fine-tuning. Experiments on the LongBench and LongBench v2 benchmarks demonstrate consistent performance improvements, with gains of up to 5% over strong baselines. Our method is model-agnostic, effective across diverse input lengths and task difficulties, and compatible with a wide range of SLLMs, including LLaMA and GLM. The proposed approach offers a practical, generalizable solution for deploying robust QA systems in resource-constrained environments.

Read the paper · More papers on PaperTik