Large language models know a lot about the world, but nothing about your return policy, internal procedures or last year's technical specifications. RAG (retrieval-augmented generation) is the most widely used way to close that gap. This article explains what RAG is, how it works and how to build an AI assistant that answers questions from your company documents.
What is RAG and why do you need it?
RAG means retrieving relevant documents before the model answers. When a user asks a question, the system searches your content, selects the most relevant passages and sends them to the model along with the question. The model then answers based on those sources.
- Freshness: update a document and the answers update, with no retraining.
- Verifiability: the assistant can cite which document an answer came from.
- Fewer hallucinations: answers are grounded in provided text. The risk is reduced, not eliminated, so design and testing still matter.
RAG vs. fine-tuning
| RAG | Fine-tuning | |
|---|---|---|
| Purpose | Give the model current, company-specific knowledge | Change style, format or behavior on a task |
| Updating data | Add or remove documents | Requires a new training run |
| Citations | Natural | Difficult |
| Access control | Filter documents per user | Knowledge is baked into the model |
For most enterprise knowledge assistants, RAG is the right starting point. Fine-tuning, if needed, is usually layered on top.
How a RAG pipeline works
1. Ingestion and cleaning
Sources can include PDFs, Word and Excel files, intranet pages, support tickets or product records. Poor scans, repeated headers and outdated versions are cleaned up here; much of the final quality is decided at this step.
2. Chunking
Documents are split into meaningful passages. Too small and context is lost; too large and irrelevant text reaches the model. Splitting along headings usually works well.
3. Embeddings and a vector database
Each passage is converted into a vector that represents its meaning and stored in a vector database, so a question about "sending an item back" can match a paragraph titled "returns".
4. Retrieval and reranking
Hybrid search, combining semantic and keyword search, performs better on product codes and proper names. A reranking step can filter the results further.
Choose the embedding model carefully if your documents are not in English. Morphologically rich languages such as Turkish or German benefit from testing several models against your own queries before committing.
5. Grounded answer with citations
The model is instructed to rely only on the provided passages, admit uncertainty and cite sources.
Enterprise use cases
- Internal knowledge base for HR policies and procedures.
- Technical support based on manuals and service notes.
- Sales teams accessing catalogs and pricing in the field, for example inside SFA software.
- Searching and comparing clauses across contracts.
- Customer-facing chatbots answering policy questions.
Access control, security and privacy
The most critical and most often skipped topic is authorization. The assistant must never answer from a document the user could not open themselves. Each passage is tagged with permissions from the source system and filtered at query time. Personal data in documents, provider retention policies and data residency should be reviewed under GDPR or KVKK where applicable.
Measuring RAG quality
- Build an evaluation set from real questions with known correct answers.
- Measure retrieval (was the right document found?) and generation (is the answer correct and faithful?) separately.
- Re-run the set after every document or prompt change.
- Collect user feedback in production.
Separating the two tells you where to focus: retrieval settings when the right document is missing, prompts or model choice when the document is found but the answer is wrong.
How to start a RAG project
Start with one department and a limited document set, such as customer service policies or service manuals. A pilot reveals data quality issues early and lets you decide on expansion with evidence. The interface can be a web panel, your intranet or a mobile app, built with our web development team. After the pilot, expansion usually means adding sources, opening access to more user groups and eventually letting the assistant take actions, which turns it into an AI agent with its own permission design. Automatic synchronization with source systems keeps the index fresh with little manual effort. For broader context, see our guide to AI integration for businesses.
BernSoftware designs and builds RAG-based assistants that work securely with your company documents. Explore our AI solutions or contact us to discuss your project.
Frequently asked questions
Which file formats can a RAG system use?
PDF, Word, Excel, PowerPoint, web pages, database records and support ticket text. Scanned PDFs need an OCR step first.
What happens when the answer is not in the documents?
A well-designed assistant says it could not find the information and points the user to the right team, enforced through instructions and testing.
Do we need to retrain the model when documents change?
No. Updated documents are re-indexed and the assistant immediately answers with the new information.
Planning a project like this?
Plan it in 10 steps