What is RAG?
Problem What is RAG (Retrieval-Augmented Generation)?
Be ready to discuss
- The core loop: embed the query, retrieve relevant passages from an external knowledge base via vector similarity search, inject them into the prompt as context, then generate an answer grounded in what was fetched.
- Why it helps: it grounds output in retrieved facts, cuts hallucination, and lets the model answer over private or fast-changing knowledge it never saw during training — without retraining the model.
- RAG vs. fine-tuning: retrieval updates knowledge by re-indexing documents rather than touching weights; fine-tuning teaches behaviour and format, retrieval supplies facts.
- The indexing pipeline: chunking strategy and chunk size, embedding model choice, and the vector store plus its ANN index (HNSW/IVF).
- Retrieval quality is usually the bottleneck, not the LLM: hybrid dense + BM25 keyword search, re-ranking the top-k with a cross-encoder, and how many chunks to fit into context.
- Failure modes: retrieving nothing relevant, retrieving contradictory passages, the model ignoring context and answering from parametric memory, and context-window/cost limits.
- Evaluation: retrieval recall@k measured separately from answer faithfulness/groundedness, plus citation coverage.
asked …