RAG and knowledge
RAG and knowledge
Learn retrieval-augmented generation, indexing, retrieval, ranking, grounding, and freshness.
flowchart LR D[Documents] --> C[Chunks] C --> E[Embeddings] E --> V[(Index)] Q[Query] --> QE[Query representation] QE --> V V --> F[Permission filter + rerank] F --> X[Grounded context] X --> M[Model] M --> A[Answer + sources]
Lesson overview
RAG and knowledge
Retrieval-augmented generation (RAG) supplies relevant external information to a model at run time. It is useful for private, changing or domain-specific knowledge that should not be assumed to live in model weights.
Ingestion
Clean documents, split them into meaningful chunks, attach metadata and create searchable representations such as embeddings. Preserve source, tenant, version, permission and update metadata.
Chunking
Very small chunks lose context; very large chunks reduce retrieval precision and consume the model context budget. Chunk by semantic boundaries where possible and test retrieval quality rather than choosing a universal size.
Query path
Transform the user's question into a search request, retrieve candidates, apply authorization filters, optionally rerank them and assemble a compact evidence set. Permission filtering must happen before content reaches the model.
Grounding
Preserve source IDs. For important answers, require citations or evidence references. Evaluate retrieval quality separately from answer quality so you can identify whether a failure came from search or generation.
When not to use RAG
If information is small and already available in context, retrieval may add unnecessary complexity. If the task requires transactional truth, query the authoritative application service instead of a stale document index.
The goal is not “use a vector database.” The goal is to supply the right evidence at the right time.
Learning path
Theory → Example → Code → Practice → Quiz → Challenge → Completion
Step 1
Theory
RAG and knowledge
Retrieval-augmented generation (RAG) supplies relevant external information to a model at run time. It is useful for private, changing or domain-specific knowledge that should not be assumed to live in model weights.
Ingestion
Clean documents, split them into meaningful chunks, attach metadata and create searchable representations such as embeddings. Preserve source, tenant, version, permission and update metadata.
Chunking
Very small chunks lose context; very large chunks reduce retrieval precision and consume the model context budget. Chunk by semantic boundaries where possible and test retrieval quality rather than choosing a universal size.
Query path
Transform the user's question into a search request, retrieve candidates, apply authorization filters, optionally rerank them and assemble a compact evidence set. Permission filtering must happen before content reaches the model.
Grounding
Preserve source IDs. For important answers, require citations or evidence references. Evaluate retrieval quality separately from answer quality so you can identify whether a failure came from search or generation.
When not to use RAG
If information is small and already available in context, retrieval may add unnecessary complexity. If the task requires transactional truth, query the authoritative application service instead of a stale document index.
The goal is not “use a vector database.” The goal is to supply the right evidence at the right time.
Step 2
Example
Example: SaaS documentation assistant
Index documentation with product, version, tenant and source metadata. A query first establishes the user's access, retrieves candidate chunks, filters unauthorized documents, reranks the candidates and sends only the best evidence to the model with source IDs.
Step 3
Code
Permission-aware retrieval
const candidates = await vector.search({ embedding, filter: { tenantId } });
const allowed = candidates.filter(c => policy.canReadDocument(user, c.documentId));
const context = allowed.slice(0, 6).map(c => ({ text: c.text, sourceId: c.documentId }));Retrieval must not become an authorization bypass.
Step 4
Practice
Practice
Design a RAG pipeline for a SaaS documentation assistant. Decide document boundaries, chunking, metadata, retrieval count, reranking, permission filtering, citation format and stale-document handling.
Step 5
Quiz
1. What is RAG primarily for?
2. When should document permissions be applied?
Step 6
Challenge
Challenge
Create 20 evaluation questions where each correct answer requires a specific source document. Measure retrieval recall separately from final answer correctness.
Complete every stage
Work through every step in order, then the lesson will be marked complete.
Each chapter and subtopic has its own public URL under /ai-agent.