Offices in Noida · Ranchi, India admin@twaratechnologies.comCareers

AI

Retrieval-augmented generation (RAG) explained: grounding AI in your own data

How retrieval-augmented generation works, from chunking and embeddings to hybrid search and re-ranking, plus access control, evaluation and the security risks.

By Twara TechnologiesPublished 8 min read

The problem RAG solves

A large language model (LLM) answers from what it absorbed during training. That knowledge stops at a cut-off date, it does not include your internal policies, product manuals or contracts, and the model cannot show you where an answer came from. Ask it about last month’s price list and it will either decline or, worse, produce something plausible and wrong.

Retrieval-augmented generation (RAG) addresses this by fetching relevant passages from a source you control and placing them in the prompt alongside the user’s question. The model then answers from that supplied context rather than from memory alone. It is a common pattern behind internal knowledge assistants, customer-support bots and document question-answering tools, because it lets a general-purpose model work with private, current information without retraining it.

Where the idea comes from

The term was introduced in a 2020 paper by Patrick Lewis and colleagues, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, accepted at NeurIPS 2020. The authors described a model’s trained weights as “parametric” memory and an external, searchable index as “non-parametric” memory. Their system paired a pre-trained sequence-to-sequence model with a dense vector index of Wikipedia, reached through a neural retriever. They reported state-of-the-art results on three open-domain question-answering tasks, and found that generated text was more specific, diverse and factual than a comparable model without retrieval.

They also named the weaknesses RAG is meant to fix: models that rely only on their weights struggle to provide provenance for their answers and are hard to update. Those two points remain the main business case today.

How a RAG pipeline works

A production RAG system has two halves: an ingestion pipeline that prepares your content, and a query pipeline that runs every time someone asks a question.

Ingestion: preparing the knowledge base

  1. Collect sources. Documents, wiki pages, tickets, database records or API responses. AWS describes this material as “external data”, meaning data that sits outside the model’s training set.
  2. Extract and clean. Pull text out of PDFs, scanned images and office files, and strip navigation, boilerplate and duplicated content.
  3. Chunk. Split long documents into smaller passages so that each one can be matched on its own. Microsoft’s Azure AI Search guidance treats chunking as a core preparation step, because returning a whole manual wastes the model’s limited context window.
  4. Embed. Pass each chunk through an embedding model, which converts text into a numeric vector that captures its meaning. Similar passages end up close together in vector space.
  5. Index. Store the vectors, the original text and useful metadata (source, date, owner, access permissions) in a vector database or a search engine with vector support.
  6. Keep it fresh. Re-index when documents change. AWS notes that documents and their embeddings need refreshing through automated real-time processes or periodic batch processing. Without that, the system can confidently quote outdated material.

Query: answering a question

  1. The user’s question is converted into a vector with the same embedding model.
  2. The retriever finds the chunks whose vectors are closest to the question, often combined with a keyword search.
  3. The best chunks are optionally re-ranked so the most useful ones come first.
  4. The application builds a prompt containing instructions, the retrieved chunks and the question.
  5. The LLM generates an answer, ideally with citations pointing back to the source chunks.

Retrieval quality decides answer quality

If retrieval returns the wrong passages, even the best model will give a poor answer. Much of the engineering effort in a RAG project therefore goes into retrieval, not the prompt.

Chunking choices

There is no universal chunk size. Small chunks match precisely but can lose surrounding context; large chunks keep context but dilute relevance and use more tokens. Splitting on natural boundaries (headings, sections, table rows) usually works better than fixed character counts. Adding a short header to each chunk with the document title and section path helps both retrieval and citation.

Pure vector search is good at matching meaning but can miss exact terms such as part numbers, error codes or names. Keyword search is the reverse. Microsoft recommends hybrid queries that run keyword and vector search in parallel “for maximum recall”, then merge the results.

Azure AI Search merges them using Reciprocal Rank Fusion (RRF). Each document receives a score of 1/(rank + k) from each result list it appears in, and the scores are summed. Documents that rank highly in several lists rise to the top. Microsoft’s documentation says the algorithm performs best with a small constant, such as 60.

Re-ranking

A re-ranker is a second model that reads the question and each candidate passage together and re-scores them by relevance. In Azure AI Search, semantic ranking runs after RRF merging. Re-ranking adds some latency, but it can noticeably improve which passages reach the model.

Metadata filters

Filters on date, product, region or document type narrow the search before similarity is computed. They are cheap, predictable and easy to explain to users.

Access control is not optional

Indexing private documents into a shared store creates a new route to data that some users should never see. Microsoft lists security and governance among the core RAG challenges, and its guidance describes document-level “security trimming” so users only retrieve content they are authorised to read.

The OWASP Top 10 for LLM Applications 2026 is more specific. Under Sensitive Information Disclosure (LLM02:2026) it advises enforcing document- and chunk-level authorisation inside the index query, not as a filter applied after retrieval. Under Vector and Embedding Weaknesses (LLM09:2026) it adds that tenant scoping must be validated server-side, because a client-supplied scope “is a suggestion, not a control”, and that a mostly public document can still contain one confidential paragraph.

In practice that means:

  • Carry each source’s permissions into the index as metadata at ingestion time.
  • Filter by the signed-in user’s identity on every query, on the server.
  • Use separate indexes for tenants or trust levels where the sensitivity justifies it.
  • Re-sync permissions when they change in the source system.

Security risks specific to RAG

RAG changes an application’s attack surface. The OWASP 2026 list highlights several risks that apply directly:

  • Indirect prompt injection (LLM01:2026). A retrieved document can contain text written to look like instructions. OWASP notes that LLMs make no architectural distinction between instructions and data, and that an injection written into a RAG corpus or vector store can affect every later session that reads from it.
  • Poisoning (LLM05:2026). Anyone who can add content to the corpus can influence answers. OWASP recommends trust boundaries, filtering of retrieved content and source scoring for RAG systems.
  • Vector and embedding weaknesses (LLM09:2026). OWASP recommends normalising content before embedding, including stripping zero-width characters, white-on-white text and look-alike Unicode characters, and recording provenance for every embedding so a compromised batch can be found and removed.
  • Misinformation (LLM07:2026). Grounding reduces errors but does not remove them. Stale, incomplete or contradictory sources still produce wrong answers delivered with confidence.

Treat every retrieved passage as untrusted input. Keep tool permissions narrow, validate any structured output in application code, and never let retrieved text alone trigger an action with real-world consequences.

Measuring whether it works

A RAG system needs its own test suite, built before launch and kept up to date.

  • Golden question set. Collect real questions with known correct answers and the source passages that support them. Include questions the system should refuse.
  • Retrieval metrics. For each question, check whether the right passage appeared in the top results. If it did not, no amount of prompt tuning will help.
  • Answer metrics. Check that answers are faithful to the retrieved text (no unsupported claims), relevant to the question and correctly cited. A mix of automated scoring and human review works best.
  • Production feedback. Log questions, retrieved chunks and answers, with appropriate privacy controls, and give users a simple way to flag a bad answer.

Run the suite whenever you change the embedding model, chunking strategy, prompt or underlying LLM.

When RAG is the right choice, and when it is not

RAG fits well when:

  • answers must come from a defined body of documents that changes over time;
  • users need citations they can check;
  • different users are allowed to see different content;
  • you want to switch LLM providers without rebuilding the knowledge layer.

AWS describes RAG as a more cost-effective way to introduce new data to an LLM than retraining a foundation model, which is why it is usually the first option to try.

Consider something else when:

  • the task is about style, format or a narrow behaviour rather than facts, where fine-tuning or better prompting may fit better;
  • the data is structured and the question is really a query (totals, filters, joins), where translating the question into a database query is more reliable than retrieving text;
  • the whole source fits comfortably in the model’s context window and rarely changes.

A practical checklist

  • Define the questions the system must answer and the sources that hold the answers.
  • Clean and chunk content on natural boundaries; keep titles and section paths with each chunk.
  • Use hybrid search and a re-ranker; add metadata filters where users expect them.
  • Enforce permissions inside the retrieval query, on the server, for every request.
  • Treat retrieved text as untrusted; restrict what the model can trigger.
  • Show citations, and let users report wrong answers.
  • Build a golden test set and re-run it on every change.
  • Schedule re-indexing and monitor index freshness.

Sources

  1. Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv:2005.11401, NeurIPS 2020)
  2. AWS: What is Retrieval-Augmented Generation?
  3. Microsoft Learn: RAG and generative AI in Azure AI Search
  4. Microsoft Learn: Hybrid search scoring (RRF) in Azure AI Search
  5. OWASP GenAI Security Project: OWASP Top 10 for LLM Applications 2026

Facts in this article were checked against the linked sources on 9 October 2026. Rules, prices and standards change; check the source before relying on a detail. This article is general information, not legal or financial advice.

Related service: AI & machine learning

Have something you want to build or fix?

Tell us what you are trying to achieve. We will reply with questions, options and an honest view of what it would take, whether or not we are the right fit.