Question 1
A financial services company is developing a Retrieval-Augmented Generation (RAG) solution to answer questions about internal compliance documents. The documents are a mix of short policy statements (1-2 paragraphs) and long procedural guides (10-20 pages). The goal is to ensure that answers are precise and source attribution is accurate. Which data preparation strategy is most suitable for this scenario?
Answer and explanation
Correct answer: B
A recursive character text splitter is ideal for documents with varied structure. It attempts to split along semantic boundaries (paragraphs, sections) first before falling back to smaller units. This preserves the context within chunks, which is crucial for accurate retrieval from both short policies and long guides. A moderate chunk size with overlap ensures that sentences or ideas are not awkwardly split between chunks.