Optimizing Context Retrieval for RAG via Heading-Aware Chunking and Hierarchical Document Structure Integration

Pham Doan Tinh, Luong Ngoc Phuong · 2025

While LLMs perform well in general contexts, they often struggle with domain-specific or time-sensitive queries. Retrieval-augmented generation (RAG) offers a more robust approach by incorporating external knowledge. Yet, its effectiveness largely depends on how well the system understands the structure of the source documents. Current methods often overlook hierarchical features found in user guides-like directory structures and nested headings-leading to poor retrieval quality. This study proposes three preprocessing techniques: Overlap-based Heading-aware Chunking, Heading Path Augmentation, and Document Source Path Augmentation. These methods preserve document structure to enhance retrieval accuracy. Evaluated on the AWS Documentation Dataset using an LLM-as-a-Judge setup, our approach significantly improves RAG performance across two embedding models, showing the value of structured context in document-based QA.

Read the paper · More papers on PaperTik