Thinklytics

AEO Primer · 4 min read · May 2026

How RAG Works: Ingestion, Retrieval, and Generation, Component by Component

By Thinklytics Partners, Practitioner Notes

RAG has two phases. An offline ingestion pass chunks your documents and writes embeddings to a vector store. An online pass embeds the query, pulls the top-k nearest chunks, and grounds the answer on them. This is the component-level view, including where retrieval quality breaks and what production systems add to fix it.

RAG (Retrieval-Augmented Generation) is a pattern where an LLM retrieves relevant context from an external knowledge store at inference time and grounds its response on that retrieved context. It is the dominant 2026 technique for giving LLMs access to fresh, domain-specific, or enterprise-private knowledge without retraining the model.

What RAG actually is

RAG has two phases, and they run at different times

  • Ingestion. Offline. Split source documents into chunks, generate an embedding for each chunk with an embedding model, store the chunks and embeddings in a vector database. This is the indexing step.
  • Inference. Online. Embed the user query with the same model, find the top-k most similar chunks by cosine similarity or hybrid lexical and semantic search, inject them into the prompt, generate the response.

The shared idea is that the model does not need to know the answer. It needs to retrieve the right context and ground the response on it.

Source: Thinklytics AI readiness practice, 2026.

RAG has two phases:

  • Ingestion (offline): split source documents into chunks, generate embeddings for each chunk using an embedding model (OpenAI Ada, Cohere Embed, Snowflake Arctic Embed, or open-source alternatives), store the chunks and embeddings in a vector database. This is the indexing step.
  • Inference (online): embed the user query using the same embedding model, find the top-k most similar chunks via cosine similarity (or hybrid lexical + semantic search), inject the retrieved chunks into the LLM prompt as context, generate the response.

The shared idea: the model does not need to know the answer. It needs to retrieve the right context and ground the response on it.

If you are briefing a non-technical stakeholder on why any of this is worth doing, the plain-English RAG guide covers the same ground without the machinery.

What people confuse it with

Three component-level confusions

Each one points at a different part of the pipeline, which is why each one produces a different kind of failure.

What people sayWhat is actually trueWhere it bites
Semantic search and RAG are the same thing.Semantic search ranks documents and stops. RAG carries the retrieved passages into the prompt and generates an answer grounded on them.You ship a search box and wonder why it does not answer questions
The vector database is the system.It is one component. Production retrieval also needs chunking suited to the document type, metadata filtering, reranking over the top-k results, and query rewriting.The store is fast and the answers are still wrong
Retrieval quality is a model problem.It is mostly a pipeline problem. If the relevant chunk is never retrieved, nothing in the generation step can recover it.Budget goes to a bigger model instead of to chunking and reranking

Source: Thinklytics AI readiness practice, 2026.

  • "Semantic search and RAG are the same thing." Semantic search ranks documents and stops there. RAG carries the retrieved passages into the prompt and has the model generate an answer grounded on them. The retrieval half is shared, the generation half is not.
  • "The vector database is the system." It is one component. Production retrieval also needs a chunking strategy suited to the document type, metadata filtering, a reranking pass over the top-k results, and query rewriting for queries whose wording sits far from the source text.
  • "Retrieval quality is a model problem." It is mostly a pipeline problem. If the relevant chunk is never retrieved, nothing in the generation step can recover it.

Where the architecture fits

What the architecture does and does not buy you

The top four are what the pipeline is for. The bottom three are what no amount of retrieval engineering will fix.

  • Selecting from a corpus too large to fit in the prompt. Retrieval narrows the input to the passages that matter instead of sending everything.
  • Updating knowledge without retraining. The index is rebuilt when the content changes and the model is never touched.
  • Returning an identifiable source chunk. Citation is only possible because retrieval hands back something addressable.
  • Reaching private and domain-specific content. General model capability does not substitute for content the model never saw.
  • Removing hallucination. It reduces it. The model can still invent facts absent from the retrieved context, or misread what was retrieved.
  • Multi-hop reasoning across many documents. A top-k pass cannot surface every chunk the answer depends on.
  • Small corpora that fit in the prompt. The retrieval step adds latency and failure modes and buys nothing.

Retrieval quality is the gating constraint. If the relevant chunk is not retrieved, the model cannot use it.

Source: Thinklytics AI readiness practice, 2026.

The pipeline earns its complexity when:

  • The corpus is too large to inject into the prompt, so the retrieval step has to select rather than send everything.
  • The content changes on its own schedule and the index has to be rebuilt without retraining anything.
  • The answer has to cite the passage it came from, which means retrieval has to return identifiable source chunks.
  • The knowledge is private or domain-specific, so no amount of general model capability substitutes for it.

What the architecture cannot do

Retrieval does not solve:

  • Hallucination. It reduces it without removing it. The model can still invent facts that were not in the retrieved context, or misread what was retrieved.
  • Multi-hop reasoning across many documents, where a top-k pass cannot surface every chunk the answer depends on.
  • A corpus small enough to fit in the prompt, where the retrieval step adds latency and failure modes and buys nothing.

How Thinklytics works on RAG

We scope RAG engagements as data-architecture-first, not vector-database-first. Retrieval quality follows from data hygiene, chunking strategy, and metadata design. See LLM grounding data architecture.

Frequently asked questions

What is RAG in one sentence?

RAG (Retrieval-Augmented Generation) is a pattern where an LLM retrieves relevant context from an external knowledge store (typically a vector database) at inference time and grounds its response on that retrieved context, reducing hallucination and enabling the model to use knowledge that was not in its training data.

How does RAG work?

Three steps. (1) At ingestion time: split documents into chunks, generate embeddings for each chunk (using a model like OpenAI Ada or Cohere Embed), store the chunks and embeddings in a vector database. (2) At query time: embed the user query, find the top-k most similar chunks via cosine similarity. (3) Inject the retrieved chunks into the LLM prompt as context, generate the response.

What problem does RAG solve?

Three problems. LLM hallucination on facts the model does not know reliably. LLM knowledge staleness (the model's training data has a cutoff). Inability to cite specific source documents in the response. RAG addresses all three by grounding the response on retrieved source content.

Is RAG the same as fine-tuning?

No. Fine-tuning adjusts the model's weights to internalize new knowledge or behavior. RAG keeps the model frozen and injects knowledge at inference time. Fine-tuning is better for stylistic or behavioral changes; RAG is better for factual knowledge that changes frequently. Many production systems use both.

What is a vector database?

A database optimized for k-nearest-neighbor search over high-dimensional vectors (embeddings). Examples include Pinecone, Weaviate, Qdrant, Chroma, pgvector (Postgres extension), and the vector features inside Snowflake Cortex Search, Databricks Vector Search, OpenSearch, and Elastic.

What is the difference between RAG and semantic search?

Semantic search is the underlying retrieval mechanism (embedding-based similarity). RAG is the broader pattern that combines semantic search with LLM generation. You can have semantic search without RAG (just retrieval), but RAG without semantic search (or some equivalent retrieval) is just an LLM.

What are RAG's limitations?

Retrieval quality is the gating constraint. If the relevant chunk is not retrieved, the LLM cannot use it. Chunk boundaries can break context. Long documents may not fit in the model's context window even after retrieval. Multi-hop reasoning (combining facts from multiple chunks) is harder than single-document RAG. Production systems address these with hybrid search, reranking, query rewriting, and structured retrieval.

How does Thinklytics work on RAG?

We scope RAG engagements as data-architecture-first, not vector-database-first. The retrieval quality follows from data hygiene, chunking strategy, and metadata design. See LLM grounding data architecture.

Topics covered

  • RAG architecture
  • retrieval pipeline
  • chunking
  • embeddings
  • vector search
  • reranking
  • semantic search
  • LLM grounding

Frequently asked questions

What is RAG in one sentence?

RAG (Retrieval-Augmented Generation) is a pattern where an LLM retrieves relevant context from an external knowledge store (typically a vector database) at inference time and grounds its response on that retrieved context, reducing hallucination and enabling the model to use knowledge that was not in its training data.

How does RAG work?

Three steps. (1) At ingestion time: split documents into chunks, generate embeddings for each chunk (using a model like OpenAI Ada or Cohere Embed), store the chunks and embeddings in a vector database. (2) At query time: embed the user query, find the top-k most similar chunks via cosine similarity. (3) Inject the retrieved chunks into the LLM prompt as context, generate the response.

What problem does RAG solve?

Three problems. LLM hallucination on facts the model does not know reliably. LLM knowledge staleness (the model's training data has a cutoff). Inability to cite specific source documents in the response. RAG addresses all three by grounding the response on retrieved source content.

Is RAG the same as fine-tuning?

No. Fine-tuning adjusts the model's weights to internalize new knowledge or behavior. RAG keeps the model frozen and injects knowledge at inference time. Fine-tuning is better for stylistic or behavioral changes; RAG is better for factual knowledge that changes frequently. Many production systems use both.

What is a vector database?

A database optimized for k-nearest-neighbor search over high-dimensional vectors (embeddings). Examples include Pinecone, Weaviate, Qdrant, Chroma, pgvector (Postgres extension), and the vector features inside Snowflake Cortex Search, Databricks Vector Search, OpenSearch, and Elastic.

What is the difference between RAG and semantic search?

Semantic search is the underlying retrieval mechanism (embedding-based similarity). RAG is the broader pattern that combines semantic search with LLM generation. You can have semantic search without RAG (just retrieval), but RAG without semantic search (or some equivalent retrieval) is just an LLM.

What are RAG's limitations?

Retrieval quality is the gating constraint. If the relevant chunk is not retrieved, the LLM cannot use it. Chunk boundaries can break context. Long documents may not fit in the model's context window even after retrieval. Multi-hop reasoning (combining facts from multiple chunks) is harder than single-document RAG. Production systems address these with hybrid search, reranking, query rewriting, and structured retrieval.

How does Thinklytics work on RAG?

We scope RAG engagements as data-architecture-first, not vector-database-first. The retrieval quality follows from data hygiene, chunking strategy, and metadata design. See [LLM grounding data architecture](/insights/llm-grounding-data-architecture).

Related reading

If this is the problem you have