Why it matters
A language model’s parameters are fixed at training time. Anything the model needs to know that is newer than its training data, specific to one organisation, or too detailed to have been memorised reliably has to arrive some other way. Retrieval is that other way: the relevant text is found first and placed in the model’s context, and the model generates from it.
The practical consequence is that a RAG system’s failures split into two kinds that are often confused. If the retriever returns the wrong documents, no amount of prompt work will fix the answer. If the retriever returns the right documents and the answer is still wrong, the problem is in generation. Systems that do not measure the two separately tend to tune the wrong one.
Origin of the term
The term was introduced by Lewis et al. in 2020, in work that combined pre-trained parametric memory — a sequence-to-sequence model — with non-parametric memory in the form of a dense vector index of Wikipedia accessed through a pre-trained neural retriever.1
Contemporary systems described as RAG frequently differ substantially from that original formulation: many use keyword or hybrid retrieval rather than a dense neural retriever, and most do not train the retriever and generator jointly. The name has become a description of the general pattern rather than of the specific architecture in the paper.
What it does not solve
Retrieval grounds a response in a document. It does not verify that the document is correct, that it is current, or that the model’s summary of it is faithful. A RAG system built over an outdated internal wiki will reproduce the wiki’s errors with more fluency than before.
Footnotes
-
Lewis et al. 2020 — full source details. ↩
Sources
The factual claims on this page are backed by the following sources.

