RAG Architecture and Overview
- What is RAG?
- Look up the most relevant passages, hand them to the model with the question, and generate the answer from that context.
- What does it fix?
- Hallucinated facts, the cost of stuffing whole documents into every prompt, and the gap between the training cutoff and private or newer files.
- What comes after native RAG?
- GraphRAG, when the question spans many documents, and Agentic RAG, when the model needs to search, retry, or call a tool.
- Where is the code?
- The companion guide builds the same pipeline in pure Python, then LangChain, FAISS, and Streamlit.
What is RAG?
Large language models have taken the world by storm, but many users encounter two common pain points: Why do models sometimes hallucinate convincing falsehoods? And why are they completely unaware of a company’s internal guidelines?
To tackle these challenges, RAG (retrieval-augmented generation) was introduced by Meta AI (Facebook AI) in their 2020 paper, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.
Think of a standalone LLM as a student taking a closed-book exam. While the student knows a vast amount of general knowledge, asking them about yesterday’s news or a private corporate budget will force them to guess. RAG acts as an open-book reference guide provided during the test.
Before responding, the system retrieves the most relevant pages from an external knowledge base and hands them to the model alongside the user’s question. The model then answers based on this context, ensuring accuracy and reliability. Thanks to its low entry barrier, fast implementation, and high ceiling, RAG has become the default architecture for enterprise knowledge bases and AI agents.
The weights still come from pretraining and later teaching. RAG does not rewrite them. It supplies the pages the model was never trained on.
3 core LLM bottlenecks solved by RAG
Standard LLMs face major hurdles in commercial deployment, all of which RAG addresses effectively.
1. Model hallucination
- Issue: Models frequently fabricate facts, citations, or specs with unearned confidence.
- Cause: LLMs operate as probabilistic next-token predictors without built-in fact-checking mechanisms.
- RAG’s solution: Supplies verifiable factual context and grounds responses strictly within the provided materials.
2. Context window and cost constraints
- Issue: While modern models (like Gemini 2.5 Pro) support millions of tokens, feeding entire multi-hundred-page documents into every prompt is slow and cost-prohibitive.
- RAG’s solution: Pinpoints and extracts only the relevant snippets (a few hundred words), cutting compute costs while preventing key details from getting lost in long contexts.
3. Knowledge cuts and missing domain expertise
- Issue: Model training data has a strict knowledge cutoff date and lacks visibility into private enterprise files (for example, internal HR policies, or Notion and DingTalk docs).
- RAG’s solution: Bypasses costly fine-tuning by dynamically connecting external data sources, instantly extending the model’s knowledge boundaries. If you do need the style of the answers to change, that is a separate step, covered in supervised fine-tuning.
Standard native RAG workflow
A classic RAG pipeline operates across two main stages: data ingestion and query processing.
[Ingestion]
Raw Docs → Text Splitting → Vector Embedding → Vector Database Storage
[Querying]
User Question → Question Embedding → Similarity Search → Prompt Construction → LLM Response
- Document parsing and splitting. Large files are chunked into smaller passages (for example, 300–500 words each) for precise matching.
- Text embedding. An embedding model converts text chunks into high-dimensional numerical vectors. Semantically similar sentences sit close to each other in vector space.
- Vector search. User queries are converted into vectors, and cosine similarity is used to pull the top-K most relevant chunks from the database.
- Prompt construction. The retrieved snippets and the original prompt are combined into a single contextual prompt (for example, “Based on the following context: [Snippets], answer this question: [Question]”).
- LLM generation. The model reads the retrieved context and generates a grounded, accurate response.
The embedding step and the cosine comparison are the whole retrieval trick. The from-scratch Python guide implements both without a framework, then rebuilds the same flow with LangChain.
Advanced RAG paradigms
As business requirements grow more complex, standard native RAG has evolved into more intelligent, reasoning-capable architectures.
1. GraphRAG (knowledge graph-based RAG)
- Background: Standard vector search relies on semantic similarity and struggles with multi-document reasoning (for example, “Summarize all supply chain risks across company projects over the past 3 years”).
- Mechanism: Extracts entities (people, companies, products) and relationships to build a knowledge graph. Retrieval combines vector search with graph traversal.
- Advantage: Enables multi-hop logic reasoning across multiple documents with higher global clarity and structure.
2. Agentic RAG
- Background: Traditional RAG follows a rigid, one-way pipeline (retrieve, then augment, then generate) that fails on complex, multi-turn inquiries.
- Mechanism: Incorporates AI agents with planning and decision-making capabilities. The model autonomously decides whether to retrieve, what keywords to query, whether to retry a search, or when to call external APIs.
- Advantage: Supports dynamic, iterative retrieval and self-correction, significantly boosting adaptability in complex workflows.
The RAG tech stack and ecosystem
Below is a breakdown of popular tools used across the RAG ecosystem.
| Category | Mainstream tools | Key features and best for |
|---|---|---|
| Out-of-the-box platforms | RAGFlow | Advanced document parsing and complex layout and table recognition; great for PDFs. |
| MaxKB | Lightweight enterprise knowledge-base assistant with no-code workflow orchestration. | |
| LangChain-Chatchat | Most popular open-source solution in China for local private knowledge bases. | |
| Frameworks and SDKs | LangChain / LangGraph | De facto standard frameworks for building production RAG and multi-agent applications. |
| OpenAI Agent SDK | Managed cloud File Search for rapid MVP validation. | |
| Vector databases | FAISS / Chroma / Qdrant | High-performance vector storage and indexing for sub-second similarity searches. |
FAISS, Chroma, and Qdrant store the vectors. The language model that reads the retrieved chunks can still be a local one. Hardware fit for that model is what the RunLocalModel checker estimates.
Related guides on this site
- Building a Python RAG System from Scratch — the implementation companion
- How LLMs Are Trained
- Supervised Fine-Tuning Explained — when you need the answers to change shape, not just the sources
- Best Local AI Models by Use Case