icaoberg / What is a RAG?

Created Fri, 10 Jul 2026 00:00:00 +0000 Modified Sat, 08 Aug 2026 21:55:55 -0400

RAG architecture Figure generated using OpenAI.

Definition

Retrieval-Augmented Generation (RAG) is a natural language processing architecture that combines a retrieval component with a generative language model. Formally, given a user query q, a RAG system first retrieves a set of relevant documents D = {d₁, d₂, …, dₖ} from an external knowledge source (such as a vector database or a document corpus), then conditions a generative model G on both the query and the retrieved documents to produce a response:

response = G(q, D)

The retrieval step typically uses dense vector search—encoding both the query and candidate documents into a shared embedding space and selecting the k nearest neighbors—though sparse methods such as BM25 are also common. The generative model then synthesizes an answer that is grounded in the retrieved evidence rather than relying solely on knowledge encoded in its parameters during training.

The architecture was formally introduced by Lewis et al. in their NeurIPS 2020 paper [1].

TL;DR

Imagine you have a really smart friend who is great at explaining things, but they stopped reading the news three years ago. That is basically a standard AI language model — knowledgeable, but stuck in the past.

Now imagine that same friend, except before answering your question they quickly Google it, skim the top results, and then explain things to you using what they just read. That is RAG.

In plain terms: instead of the AI trying to remember everything from training, it first searches a pile of documents for relevant information, then uses that information to write you a response. Think of it as giving the AI a cheat sheet that it looks up on the fly.

The two moving parts are simple:

  1. Retriever — finds the most relevant chunks of text from your documents (like a search engine).
  2. Generator — reads those chunks and writes a human-friendly answer (like a tutor summarizing what they just read).

That is it. Search first, then answer.


Why RAG instead of training from scratch?

Training a large language model from scratch is resource-intensive and inflexible. RAG offers several compelling advantages:

  1. Cost efficiency
  2. Up-to-date knowledge
  3. Source attribution and transparency
  4. Domain adaptation without fine-tuning
  5. Smaller model footprint
  6. Easier knowledge updates and corrections

Summary

RAG separates what a model knows from how a model reasons, storing facts externally and retrieving them on demand. This makes it a practical, cost-effective, and maintainable alternative to training monolithic models that must encode all world knowledge in their parameters. For most knowledge-intensive applications—question answering, document search, enterprise chat—RAG is the right first choice before committing to the far greater expense of training from scratch.


References

[1] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems, vol. 33, pp. 9459–9474, 2020. arXiv:2005.11401