What is Retrieval-Augmented Generation (RAG)?
Retrieval-Augmented Generation represents a paradigm shift in how large language models (LLMs) access and utilize information. Unlike traditional LLMs that rely solely on weights learned during training, RAG systems augment the generation process by retrieving relevant documents from an external knowledge base before producing responses. This architecture consists of three core components: a retriever, a vector database, and a generator. The retriever identifies semantically similar documents using vector embeddings, the vector database stores and indexes these embeddings for fast lookup, and the generator produces contextually grounded responses using both the retrieved documents and the original query.
Vector Space Mechanics: The Foundation of RAG
At the heart of every RAG system lies the vector spaceâa high-dimensional mathematical space where text documents and queries are represented as numerical vectors (embeddings). Each embedding is typically 384 to 1536 dimensions, depending on the embedding model used (e.g., OpenAI's text-embedding-3-small produces 1536-dimensional vectors, while sentence-transformers/all-MiniLM-L6-v2 produces 384-dimensional vectors).
The fundamental operation in RAG vector spaces is similarity computation, most commonly using cosine similarity. Cosine similarity measures the angle between two vectors, ranging from -1 to 1, where 1 indicates identical direction (perfect similarity) and -1 indicates opposite directions. The formula is: cosine_similarity(A, B) = (A · B) / (||A|| à ||B||). In practice, normalized vectors make this computation equivalent to the dot product, enabling efficient batch operations.
Production RAG System Architecture
In production systems, RAG typically operates as follows: (1) A user query arrives and is immediately embedded using the same embedding model that indexed the knowledge base; (2) The query embedding is compared against all document embeddings in the vector database using cosine similarity; (3) The top-k documents (usually 3-10) with highest similarity scores are retrieved; (4) These documents are formatted into a prompt context and sent to an LLM generator; (5) The generator produces a response grounded in the retrieved context.
The vector database layer is critical for production performance. Systems like Pinecone, Weaviate, Milvus, and Qdrant use approximate nearest neighbor (ANN) search algorithmsâsuch as hierarchical navigable small worlds (HNSW) or product quantizationâto retrieve top-k results in milliseconds rather than seconds. These indices trade perfect accuracy for speed, meaning the retrieved documents are approximately (not exactly) the most similar.
Embedding Model Selection and Its Implications
The choice of embedding model fundamentally shapes the vector space geometry. Different models produce embeddings with different properties: some cluster semantically similar content tightly, others spread it across the space. Popular production models include OpenAI's text-embedding models (trained on massive web data), Anthropic's models, open-source sentence-transformers, and domain-specific embedders fine-tuned for specialized vocabularies (medical, legal, financial).
Critically, the embedding model becomes a permanent fixture of your production system. Once documents are embedded and indexed, changing models requires re-embedding the entire knowledge baseâa process that can take hours or days for large repositories. This creates organizational inertia around embedding choices.
Vector Space Degradation Scenarios
In production, vector spaces degrade through several mechanisms. First, knowledge base drift occurs when documents are added, modified, or deleted without corresponding updates to retrieval logic. Second, query distribution shift happens when user questions diverge from patterns the embedding model was trained on. Third, semantic drift emerges when terminology or meaning evolves (e.g., "cloud" shifting from weather to computing). Fourth, model obsolescence occurs when newer, superior embedding models become available but the organization cannot justify re-embedding costs.
These degradation modes manifest as retrieval failures: queries that should match relevant documents no longer do, top-k results become irrelevant, and downstream LLM responses become hallucinated or off-topic. Monitoring these failures requires systematic tracking of embedding distribution propertiesâthe focus of this masterclass.