LongAttn: Selecting Long-context Training Data via Token-level Attention
Longyun Wu, Dawei Zhu, Guangxiang Zhao, Zhuocheng Yu, Junfeng Ran, Xiangyu Wong, Lin Sun, Sujian Li · 2025
With the development of large language models (LLMs), there has been an increasing need for significant advancements in handling long contexts.To enhance long-context capabilities, constructing high-quality training data with long-range dependencies is crucial.Existing methods to select long-context data often rely on sentence-level analysis, which can be greatly optimized in both performance and efficiency.In this paper, we propose a novel tokenlevel framework, LongAttn, which leverages the self-attention mechanism of LLMs to measure the long-range dependencies for the data.By calculating token-level dependency strength and distribution uniformity of token scores, LongAttn effectively quantifies long-range dependencies, enabling more accurate and efficient data selection.We filter LongABC-32K from open-source long-context datasets (ArXiv, Book, and Code).Through our comprehensive experiments, LongAttn has demonstrated its excellent effectiveness, scalability, and efficiency.We have released our code and the highquality long-context training data LongABC-32K at https://github.com/Lyun0912-wu/ LongAttn.