What is RAG, and what can it do with your company’s documents?
By the E-Solutions Web editorial team. Published , updated . How we write
The short answer
RAG, short for retrieval-augmented generation, is a way of making an AI answer from a chosen set of documents instead of from memory alone. Before it writes, the system searches your files for the passages that match the question, hands them to the language model and asks it to answer from them, with the source. The quality of the answer then depends on your documents.
What is RAG? The definition in one paragraph
Retrieval-augmented generation is a design in which an AI looks things up before it answers. A language model on its own writes from what it absorbed during training. That knowledge stops at a date, it never included your contracts or procedures, and the model cannot tell you where a given sentence came from. RAG adds a search step in front of the model: find the relevant passages in a collection you choose, then write the answer from them.
The term comes from a research paper first posted in May 2020 and accepted at the NeurIPS conference that year. Its authors named the problem plainly: for language models, “providing provenance for their decisions and updating their world knowledge remain open research problems.” Their answer was to pair the model with an external, searchable memory. Companies use the same pairing today to put AI to work on their own documents, because it answers both problems at once: updating a fact means replacing a document, and every answer can point to its source.
For a decision-maker, the useful mental picture is a new colleague with an excellent memory for language and no knowledge of your company, sitting next to a well-organized archive. RAG is the rule that makes them check the archive before they speak.
How a RAG assistant answers a question
Every RAG system runs the same four steps, and the quality of each one decides the quality of the answer. The first happens before anyone asks anything; the other three happen in the seconds after a question.
- Prepare the documents. Files are collected from where they live (shared drives, intranet, ticketing tool, product database), converted to text, split into passages of a few paragraphs, and each passage is turned into a vector: a list of numbers that captures its meaning. Those vectors go into an index.
- Search by meaning. The question is converted the same way and compared with every passage. The closest ones come back, even when they use different words. “How much can I spend on a hotel room” finds the passage titled “accommodation limits”.
- Answer from the passages. The model receives the question and the retrieved passages, with an instruction to answer only from them and to say so when they do not contain the answer.
- Show the sources. Each claim in the answer points to the passage it came from, so the reader can open the original and check.
You can watch those four steps in our live knowledge assistant demo. It has read the eight internal procedures of a fictional distributor, and each answer lists the numbered passages it used. Ask it something the procedures do not cover and see how it responds: that behavior matters more than any fluent answer.
Why not skip the search and paste every document into the prompt? Models now accept very long inputs, but a 2023 study found that their performance “significantly degrades when models must access relevant information in the middle of long contexts.” Selecting the few passages that matter is still the more reliable route, and it keeps each request smaller.
RAG, fine-tuning or a longer prompt: which one fits
If the goal is answers grounded in your documents, RAG is usually the right starting point, and the other two techniques work best alongside it. The three are often confused in vendor proposals, so it helps to see what each one actually changes.
| Criterion | RAG | Fine-tuning | Everything in the prompt |
|---|---|---|---|
| What it changes | What the model reads before answering | How the model writes and behaves | What the model reads, all at once |
| Where the facts live | In your documents, outside the model | Blended into the model’s weights | In each request |
| Updating a fact | Replace the document | Train again | Edit the prompt |
| Showing the source | Yes, passage by passage | No | Possible, less precise on long inputs |
| Restricting by user | Yes, by filtering the search | No | Only by building one prompt per profile |
| Good fit | Procedures, contracts, product data, support history | A house style, a fixed format, a specialized vocabulary | A handful of short, stable documents |
In practice, a solid assistant combines them: RAG for the facts, a carefully written instruction for tone and format, and, rarely, some fine-tuning when a vocabulary or output format must be followed exactly.
The three limits to check before you buy
RAG does not make an AI trustworthy by itself. Three limits decide whether an assistant helps your teams or misleads them, and a bigger model fixes none of them.
Your documents set the ceiling
An assistant that answers from your documents is only as good as those documents. If three versions of the travel policy sit in three folders, it may quote the oldest. If a key procedure exists only as a scanned PDF with no text layer, the assistant cannot read it. If the answer to a frequent question lives in someone’s head, no search will find it.
That is why a serious project starts with the corpus: pick the authoritative source for each topic, retire outdated versions, extract text from scans, and add the metadata (owner, date, audience) that lets the system prefer current documents. That work also benefits the people who search those files by hand today.
Access rights must follow the person asking
Most companies do not want every employee to read every document. An HR file, a board memo or a customer contract should only surface in answers for people who could already open it. The search step has to apply the same permissions as your document systems, for every question.
This is a real security concern. The OWASP Gen AI Security Project lists “Vector and Embedding Weaknesses” among its top ten risks for LLM applications in 2025, warning that “inadequate or misaligned access controls can lead to unauthorized access to embeddings containing sensitive information.” Its recommended fix is “permission-aware vector and embedding stores.” Ask any supplier how their design meets that recommendation, and ask to see it working with two users who have different rights.
Hallucinations shrink without disappearing
The US National Institute of Standards and Technology calls the problem “confabulation”: “the production of confidently stated but erroneous or false content.” RAG reduces it because the model has the right text in front of it. It does not eliminate it. The model can still misread a passage, combine two passages wrongly, or fill a gap with a plausible guess.
The clearest public measurement comes from law, a field where every citation can be checked. A preregistered study published in May 2024 tested commercial legal research tools whose vendors promoted RAG as a way to avoid hallucinations. It found that they “each hallucinate between 17% and 33% of the time,” fewer errors than a general-purpose chatbot, but far from none. The lesson for any company is to design for residual errors: visible sources, an assistant allowed to say “I don’t know”, and a person in the loop wherever an answer triggers a decision.
How to judge a RAG assistant before you commit
Judge it on your own questions, with the answers checked against your documents, before anyone talks about rollout. A polished demo on the supplier’s sample files tells you little about your own documents, and a short, structured test fills that gap.
- Collect 30 to 50 real questions from the people who will use it, including awkward ones: vague wording, questions spanning two documents, questions whose answer is not in the corpus at all.
- Write the expected answer and its source for each one, with the document owners. This is your reference set, and you will reuse it every time the system changes.
- Score four things per answer: is it correct, does it cite the right passage, did it refuse when it should have, and did it leak anything the tester was not entitled to see.
- Repeat the test after each change of model, document set or instruction. An assistant that was right in June can drift after an update.
Two warning signs deserve a hard stop: answers without sources, and an assistant that never says it does not know. Both mean errors will reach your teams without anyone noticing.
What a first RAG project looks like
The projects that work start narrow: one team, one well-defined document set, one place where people already ask questions. Customer support answering from product sheets and warranty terms, sales checking contract clauses, new hires asking about internal procedures. Each has a clear owner for the documents and an easy way to measure whether answers are right.
From there, the build usually follows the same path. The corpus is cleaned and connected so it stays up to date automatically. Access rights are mapped from your existing systems. The assistant is placed where people already work: a chat tool, an intranet or a support desk. Every answer is logged, so errors can be traced to a document or a retrieval problem and fixed at the source. Our AI knowledge base service covers that path end to end. When the documents are too sensitive for a public AI service, the same design runs on a private LLM hosted in Europe, and when the audience is your customers, the same retrieval powers an AI chatbot for your website that answers from your public documents.
Two regulatory points belong in the plan from day one. Personal data in the documents stays subject to the GDPR, including the rights to rectification and erasure (Articles 16 and 17), which is easier when the facts live in documents rather than in a model. And where people chat with the assistant, Article 50 of the EU AI Act requires that they be informed they are interacting with an AI system, unless that is obvious: our guide to EU AI Act compliance explains what that means for a company that deploys one. If the assistant is also meant to take actions in your tools, you are moving from RAG to an AI agent, and the design questions change.
The best way to judge all this is to try it. Ask our live knowledge assistant demo a hard question and see how it cites its sources, or a question its documents do not cover and see it say so. Nothing you type is stored. Then, if you want the same thing on your own files, book your free 30-minute assessment: we look at your document set, the access rights it needs and the questions your teams ask most, and tell you what a first assistant would take. We reply within one business day.
Frequently asked questions
What is RAG in generative AI?
It is a technique that connects a language model to a searchable collection of documents. For each question, the system first retrieves the most relevant passages, then asks the model to write its answer from those passages and to cite them. The model’s general knowledge is still there, but the facts come from your sources.
What is the difference between an LLM and RAG?
An LLM is the language model itself: it writes fluent text from what it absorbed during training, which stops at a date and never included your internal files. RAG is an architecture built around an LLM. It adds a search step over your documents, so the model answers from current, specific material it was never trained on.
Does RAG stop AI hallucinations?
It reduces them without removing them. A 2024 preregistered study of commercial legal research tools built on RAG found that they still produced hallucinated answers between 17% and 33% of the time. Showing sources, testing on real questions and letting the assistant say it does not know are what keep the residual errors visible.
Can a RAG assistant respect who is allowed to see what?
Yes, if it is designed for it. The search step must filter passages by the rights of the person asking, using the same permissions as your document systems. OWASP lists weak access control in vector databases among the main security risks of RAG applications, so this is a point to check explicitly before any deployment.
Is RAG better than fine-tuning a model?
For answering from company documents, usually yes. Fine-tuning changes how a model writes or behaves, but it is a poor way to store facts that change, and it cannot show where an answer came from. RAG keeps the facts in documents you can update, delete and restrict, and each answer can point to its passage.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv (paper accepted at NeurIPS 2020), first version 2020-05-22, accessed 2026-10-01.
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, arXiv, version of 2024-05-30, accessed 2026-10-01.
- Lost in the Middle: How Language Models Use Long Contexts, arXiv (accepted in Transactions of the Association for Computational Linguistics), first version 2023-07-06, accessed 2026-10-01.
- LLM08:2025 Vector and Embedding Weaknesses, OWASP Gen AI Security Project, accessed 2026-10-01.
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1), National Institute of Standards and Technology, July 2024, accessed 2026-10-01.
- Regulation (EU) 2016/679 (General Data Protection Regulation), EUR-Lex, Publications Office of the European Union, published 2016-05-04, accessed 2026-10-01.
- Regulation (EU) 2024/1689 (Artificial Intelligence Act), EUR-Lex, Publications Office of the European Union, published 2024-07-12, accessed 2026-10-01.