Why this decision keeps getting sold wrong
Fine-tuning sounds like ownership ("our own model!") and consultancies like it because it bills well. RAG sounds like plumbing and gets undersold. In practice, most business AI features need neither, a well-engineered prompt over a capable model solves a shocking share of use cases. This page exists so you can walk into any vendor conversation knowing which of the three you actually need, and what each will cost you to build and keep alive.
The three tools, one paragraph each
Prompt engineering, writing precise instructions and examples into the prompt itself. Zero infrastructure, zero training, changes in seconds. Ceilings: context length, and discipline at scale.
RAG (retrieval-augmented generation), the model looks up relevant passages from your documents at question time and answers from them, with citations. Answers stay current and grounded; the knowledge lives in a searchable store you update like content.
Fine-tuning, training a model further on your examples so its style, format or niche judgment shifts permanently. The knowledge is baked in; updating it means retraining.
The decision table
| Your need | Right tool | Why |
|---|---|---|
| Answer questions from your docs/handbook/policies | RAG | Knowledge changes; citations matter; hallucination must stay low |
| Consistent brand voice / strict output format at volume | Fine-tune (sometimes) | Format/style stability beats prompting at scale, after prompting genuinely fails |
| A chatbot that sounds on-brand and uses the docs | RAG + prompt engineering | Voice via instructions; facts via retrieval |
| Classification, extraction, routing of well-formed inputs | Prompt engineering first | Cheapest; fine-tune only for hard accuracy ceilings |
| "Teach it our proprietary expertise" | Usually RAG, honestly | Documents beat baked weights for anything that changes or needs sourcing |
| Anything a good prompt solves today | Prompt engineering | Every tier up costs build + monthly + maintenance |
The costs nobody quotes upfront
RAG: build is moderate (ingestion pipeline, chunking, retrieval quality tuning, the tuning is the real work), monthly is your model usage plus vector store pennies, and maintenance is content ops: stale documents produce confident wrong answers.
Fine-tuning: data preparation dominates (hundreds to thousands of cleaned examples, the dataset is the project), retraining every time knowledge or style drifts, evaluation infrastructure to prove the tune actually improved anything, and a subtle trap: fine-tuned models drift as base models deprecate. Every platform migration is a re-tune.
Prompt engineering: nearly free to build, cheap to run, and the maintenance is versioning discipline, prompts are code now, whether vendors admit it or not.
The honest escalation path
Start at prompts. Measure failures. If answers are stale or unsourced → RAG. If format/voice stability fails at volume after prompt discipline → consider fine-tuning, with an evaluation set agreed before training. Skipping steps here is how AI budgets die: fine-tuned first, tuned again each deprecation, retrieval bolted on anyway for facts, three systems where one prompt would have held.
How this practice builds it
Both, with the escalation path enforced: RAG builds with citation-first answers and content-ops handover, fine-tuning only with a measured failing baseline and a post-tune evaluation gate. If a conversation settles this in an hour instead of a retainer, it settles in an hour, that is the consulting engagement, or just ask directly.