Short definition

RAG (Retrieval-Augmented Generation) is a technique in which a large language model first retrieves content relevant to a question from a search engine, database or document collection, then bases its answer on that retrieved content.

The term was popularized by a 2020 paper from researchers at Facebook AI Research (now Meta AI). The idea is simple: the model doesn’t have to know everything by heart. It finds the relevant documents first, then writes its answer by looking at them.

How it works

  1. Retrieve: a search runs based on the user’s question. That might be the web, a company’s internal documents or a database.
  2. Augment: the passages it finds are handed to the model together with the question.
  3. Generate: the model writes its answer from those passages and often shows its sources.

ChatGPT search, Perplexity and Google’s AI Overviews can all be seen as different implementations of this approach.

Why it matters

RAG lets a model use information from after its training cutoff and lowers the risk of hallucination, though it doesn’t remove it. A model can still summarize a source incorrectly.

For GEO, the takeaway is that your site has to be findable at the retrieval step. If a page can’t be crawled, if AI crawlers are blocked, or if the key information is buried deep in the page, the model can’t use it.

Example

When someone asks Perplexity “what is the minimum wage in 2026?”, the model doesn’t know that from training. It searches the web for current sources, builds the answer from the pages it finds and lists them underneath.

Let us measure your AI visibility for free