Towards efficient multi-legal document summarization: An ensemble approach for Turkish law
Maha Ahmed Abdullah Albayati, Kürşat Mustafa Karaoğlan, Oğuz Fındık · Engineering Science and Technology an International Journal · 2025
Legal document summarization is a critical yet challenging task due to the structural complexity, terminological density, and contextual depth of legal texts, especially in low-resource languages like Turkish. Building on our previous research in hybrid summarization (Albayati and Findik, 2025), this study proposes an advanced ensemble-based framework tailored to Turkish legal documents. The framework integrates four transformer models, LED, Long-T5, BART-Large, and GPT-3.5 Turbo, and employs three complementary techniques: the Consecutive-Aware Semantic Voting Mechanism (CASVM), the Weighted Hybrid Sentence Ranking Framework (WHSRF), and a T5-Based Meta-Model (T5-Meta). Together, these components leverage both extractive precision and abstractive fluency. Experimental results over a dataset of 2,000 Turkish court decisions show significant improvements in summary quality and consistency. The proposed ensemble achieved ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-Sum scores of 0.90, 0.78, 0.88, and 0.85 respectively, reflecting relative gains of 63.64%, 122.86%, 109.52%, and 93.18% compared to the best-performing individual models. In parallel, BERTScore evaluations confirmed a high level of semantic fidelity, validating the model’s ability to preserve legal meaning even in paraphrased output. This research sets a new benchmark in Turkish legal summarization, showcasing the power of ensemble learning for multilingual legal NLP and paving the way for its integration into intelligent legal assistants that support faster, more accurate information access for legal professionals.