Most enterprise chatbot projects do not fail because the model is “not smart enough”. They fail because the system answers from a mash of marketing copy, stale PDFs and hope. Retrieval-augmented generation (RAG) is the pattern we start with: collect the source of truth, chunk it, embed it, retrieve it, generate an answer with a citation, and say “I don’t know” when the sources do not support a reply.
Fine-tuning has a place. It is for behaviour and style — how the assistant talks, how it uses tools, how it follows a house format — on top of a model that already reasons. It is a poor filing cabinet. The moment your return policy changes, a fine-tuned model is wrong until you spend another training cycle. A RAG index can be updated the same afternoon.
Our production stack is boring on purpose: an ingest pipeline, a vector index with access control, an evaluation set of real questions, guardrails that block off-topic and low-confidence answers, and a handoff to a human with the transcript attached. Citations are not decoration. They are how your support lead audits the bot on Monday.
When we do fine-tune, it is after a RAG prototype has shown the use case is real. Mixing the two without measurement is how teams burn budget on a demo that cannot be governed.
If you have a sample of documents and twenty questions your team actually answers, we can put a proof of concept on that corpus and score it. That is more useful than a workshop slide about “AI transformation”.
