Fused Transformers: Fused Information of Arabic long Article for Summarization
Zeyad Ezzat, Ayman Khalfallah, Ghadad Khoriba · Procedia Computer Science · 2024
This paper introduces a novel approach for extending the applicability of pre-trained models to accommodate longer texts. Addressing the inherent limitation of quadratic performance in attention models within transformers, we propose the concept of fused transformers. By integrating new adaptor techniques, we enhance the encoding of lengthy text segments by breaking them into shorter spans. Subsequently, these segments are fused to increase the model's effectiveness in processing extended texts. This fusion mechanism serves to fortify the model's capacity for fine-tuning. Furthermore, we implement a length-based curriculum for expedited training Our experiments yielded 16 Rouge-2 points, representing a doubling of the score achieved by the vanilla fine-tuning method on the newly introduced ”Mukhtasar” dataset for summarization. This highlights the effectiveness of managing complex relations among text segments and confirms that our method can outperform conventional training approaches.