Key takeaways
- Retrieval-augmented generation searches your own documents first, then has the AI answer using only those passages, with sources attached.
- The core pipeline is chunking, embeddings, vector search, reranking and generation, and weak chunking or retrieval causes most quality problems.
- RAG suits questions answered from policies, manuals and contracts, but numeric analysis across records is better handled by databases or BI tools.
- Clean, current, well-owned documents and permission-aware search matter more to results than which AI model you choose.
- Measure accuracy with an evaluation set of real questions, including ones with no answer, and rerun it after every change.
A general AI model knows a lot about the world, but it knows nothing about your price list, your employee handbook or the contract you signed last month. Ask it about those and it will either refuse or, worse, make up a plausible answer. Retrieval-augmented generation, usually shortened to RAG, is the most common way to fix that.
This guide explains how RAG works in plain English, what it does well and poorly, and what a typical project looks like, so you can judge whether it fits your business and ask vendors the right questions.
What RAG is, in one paragraph
Retrieval-augmented generation is a pattern where the AI system first searches your own documents for passages relevant to a question, then hands those passages to a large language model along with the question and instructions to answer using only that material. The model writes the answer, and the system shows which documents it came from.
Think of it as an open-book exam. The model does not need to memorize your documents. It needs to be given the right pages at the right moment, and it needs to stick to them.
How retrieval-augmented generation works, step by step
A RAG system has two phases: preparing your documents ahead of time, and answering questions when they arrive.
1. Chunking: splitting documents into passages
Long documents are split into smaller pieces, often a few paragraphs each. This step is called chunking. Chunks that are too small lose context; chunks that are too large bury the relevant sentence in noise. Good chunking respects the structure of the document, such as headings, sections, table rows and numbered clauses, rather than cutting every fixed number of characters.
2. Embeddings: turning text into numbers
Each chunk is converted into an embedding, which is a long list of numbers that represents its meaning. Embeddings let the system find passages that mean the same thing even when the words differ. A question about "time off for a new baby" can match a policy titled "Parental leave."
3. Vector search: finding relevant passages
The embeddings are stored in a vector database or a search index that supports vector search. When a question comes in, it is also turned into an embedding, and the system retrieves the chunks whose embeddings are closest to it. Many production systems combine this with traditional keyword search, which helps with exact terms like product codes, part numbers and names.
4. Reranking: putting the best passages first
The first search step is fast but rough. A reranking model then reviews the top candidates more carefully and reorders them by how well they actually answer the question. This step often makes a noticeable difference in answer quality, especially across large document sets.
5. Generation with sources
Finally, the top passages are placed into the model's prompt along with the question and instructions such as "answer only from the provided sources, and say so if the answer is not there." The model writes a response, and the system attaches citations or links to the source documents so users can verify the answer.
What RAG is good at, and where it struggles
RAG is a strong fit for some jobs and a poor fit for others. Knowing the difference up front saves wasted effort.
| Good fit | Poor fit |
|---|---|
| Answering questions from policies, manuals, contracts and help articles | Calculations across many records, such as "total sales by region last quarter" |
| Finding the right clause or procedure across thousands of pages | Questions that need every document reviewed, such as "list all contracts with this clause" |
| Giving answers with links back to the source | Content that changes by the minute, unless retrieval connects to live systems |
| Content that updates regularly, since you re-index rather than retrain | Messy, contradictory or outdated document sets |
For numeric and analytical questions, a database query or a business intelligence tool is usually the better path. Some systems combine both, using RAG for documents and structured queries for numbers.
Preparing your data and handling permissions
Most RAG problems trace back to the documents, not the model. Time spent on preparation pays off more than time spent on model selection.
Clean up the source content
- Remove duplicates and outdated versions. If three versions of the travel policy exist, the system may quote the wrong one.
- Fix extraction problems. Scanned PDFs need OCR, and tables, headers and footers often come out garbled. Check a sample of extracted text by eye.
- Add useful metadata. Document type, owner, department, effective date and region help with filtering and ranking.
- Assign owners. Someone should be responsible for keeping each content area current.
Respect existing access rules
A RAG system must not become a way around your permissions. If a salesperson cannot open the HR folder in SharePoint, the assistant should not quote from it to them either. The usual approach is to store access information with each chunk and filter results by the user's identity at search time. Confirm how any vendor or partner handles this before you load sensitive content, and involve your compliance team for regulated data.
Measuring accuracy and controlling hallucinations
An AI hallucination is a confident answer that is not supported by the facts. RAG reduces hallucinations because the model is working from real passages, but it does not eliminate them. Accuracy has to be measured, not assumed.
Build an evaluation set
Collect 50 to 200 real questions from the people who will use the system, along with the correct answers and the documents they come from. Include hard cases: questions with no answer in the documents, questions where two documents conflict, and questions phrased in unusual ways. Run this set every time you change chunking, models or prompts.
Measure retrieval and answers separately
- Retrieval quality: did the right passage appear in the top results?
- Answer faithfulness: does the answer stick to what the passages say?
- Answer correctness: is the final answer right and complete?
- Refusal behavior: does the system say "I don't know" when the answer is not in the documents?
Separating these tells you where to fix things. If retrieval is failing, better chunking or reranking helps. If retrieval works but answers are wrong, prompts or the model need attention.
Practical hallucination controls
- Instruct the model to answer only from the retrieved sources and to say when it cannot find an answer.
- Show citations with every answer, so users can check the source in one click.
- Set a relevance threshold so weak matches trigger a refusal rather than a guess.
- Route high-stakes topics (legal, medical, financial) to a person, or label answers as informational only.
- Review a sample of real conversations each week, especially ones users rated poorly.
Typical RAG project steps
Timelines vary with the volume and condition of your documents, but most projects follow a similar path.
- Define the use case and users. For example, "help support agents answer warranty questions" is better than "chat with all our files."
- Inventory and assess the content. Identify sources, formats, owners, sensitivity and how often each changes.
- Prepare and index a first content set. Start with one well-maintained area rather than everything at once.
- Build the evaluation set. Agree on target accuracy and refusal behavior with the business owner.
- Build and tune the pipeline. Adjust chunking, search settings, reranking and prompts against the evaluation set.
- Pilot with real users. Collect feedback and review logs.
- Expand and operate. Add sources, automate re-indexing when documents change, and keep monitoring quality.
Build vs off-the-shelf
You do not always need a custom build. Several categories of tools offer RAG-style features today:
- Built-in assistants in tools you already use. Productivity suites, help desk platforms and intranet tools increasingly include AI search over their own content. These are quick to try but usually limited to content inside that product.
- Managed knowledge assistant platforms. These connect to several sources and handle indexing for you, typically priced per user per month or by usage. Check each vendor's current pricing page and data handling terms.
- Custom or semi-custom builds. These use cloud AI services and your own retrieval pipeline. They take more effort but give you control over data sources, permissions, evaluation, integrations and user experience.
A custom approach tends to make sense when content is spread across many systems, permissions are complex, accuracy requirements are high, or the assistant must be embedded in your own application. Off-the-shelf is often the right start when content lives mostly in one platform and the use case is general.
Next steps
Before talking to vendors, pick one specific use case, list the documents it depends on and gather 30 to 50 real questions people ask today. That small package will tell any partner most of what they need to estimate effort and spot data problems early.
If you want help assessing your documents or designing an evaluation plan, a partner experienced in document AI and knowledge search can review your sources, permissions and goals and recommend whether to buy, build or combine the two. Invictus Hub offers this kind of assessment and build work. Reach out if you would like to talk it through.



