The short answer
For most business assistants that answer questions from company knowledge, start with RAG (retrieval-augmented generation). Use fine-tuning when you need the model to follow a specific format, style or narrow task very consistently. Many production systems use both.
What is RAG?
Retrieval-augmented generation keeps your knowledge outside the model. When someone asks a question, the system searches your documents, tickets or database for the most relevant passages, then gives those passages to the model along with the question. The model writes an answer based on what was retrieved, and can cite its sources.
Because the knowledge lives in a search index, updating it is as simple as updating the documents. Permissions can be enforced at search time, so people only see answers from content they are allowed to access.
What is fine-tuning?
Fine-tuning continues training a model on your own examples, so it learns a behaviour: a format, a tone, a classification scheme or a specialised task. The knowledge or behaviour becomes part of the model's weights.
Fine-tuning is excellent at teaching how to respond, but it is a poor way to teach facts that change: every update means retraining, and it is hard to show where an answer came from.
RAG vs fine-tuning, side by side
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Answering from changing company knowledge | Consistent format, tone or narrow tasks |
| Keeping knowledge current | Update documents, no retraining | Retrain for every change |
| Citing sources | Yes, answers link to documents | Not naturally |
| Access control | Enforced at retrieval time | Hard: everything is in the weights |
| Upfront effort | Lower: index content, tune retrieval | Higher: curate training data, train, evaluate |
| Response speed | Longer prompts per request | Shorter prompts, possibly a smaller model |
| Main risk | Poor retrieval leads to poor answers | Outdated or invented facts |
A simple decision guide
- Does the answer depend on facts that change (policies, prices, product docs, customer records)? Use RAG.
- Do people need to see where an answer came from? Use RAG.
- Do different users have different access rights? Use RAG with permission-aware retrieval.
- Is the problem mainly about format or style, such as always producing a specific report structure or brand voice? Consider fine-tuning.
- Is it a narrow, high-volume task such as classification or extraction, where a smaller fine-tuned model could cut latency? Consider fine-tuning.
- Is it both? Combine them: a fine-tuned model that follows your format, fed by RAG with current knowledge.
How a production RAG system works
A RAG prototype can be built in an afternoon. A RAG system people trust takes more care at each stage:
- Ingestion: documents are collected from their sources, cleaned, and split into passages that keep their headings and context, with metadata such as owner, date and access rights.
- Indexing: each passage is stored with an embedding for meaning-based search, and usually also with keywords, because exact terms like product codes matter.
- Retrieval: a question triggers a hybrid search, filtered by the user's permissions, and a re-ranking step picks the few passages that really answer it.
- Generation: the model answers using only the retrieved passages, cites them, and says when the documents do not contain the answer.
- Evaluation and monitoring: a fixed set of real questions is re-run after every change, and live answers are sampled and reviewed.
Most quality problems in RAG come from the first three stages, not from the model. That is why improving chunking, metadata and retrieval usually beats switching to a bigger model.
Common mistakes to avoid
- Fine-tuning to add knowledge. It is hard to keep current and the model may still invent details.
- Treating RAG as plug and play. Retrieval quality decides answer quality: chunking, metadata, hybrid search and re-ranking all matter.
- Skipping evaluation. Without a test set of real questions and expected answers, you cannot tell whether a change made things better or worse.
- Ignoring permissions. An assistant that can see everything will eventually show someone something they should not see.
Where to start
Pick one high-value use case, collect fifty to one hundred real questions with good answers, and build a small RAG prototype measured against that set. The results will tell you whether retrieval alone is enough or whether fine-tuning would add value. If you would like help, see our LLM and AI agent development services or start a project.