AI and LLMs · Comparison
RAG vs fine-tuning
Both make a general model fit your product. RAG hands the model the right information at question time; fine-tuning changes how the model behaves. Many teams start with RAG and fine-tune later, if at all.
2 options · 8 questions side by side · updated
| Compare | RAG | Fine-tuning |
|---|---|---|
| What it changes | What the model knows for this answer | How the model behaves every time |
| Good for | Facts, documents, policies, catalogues | Tone, format, narrow classification tasks |
| Keeping it current | Update a document; the next answer uses it | Retrain to change anything |
| What you need | Documents and a search index | Hundreds of good example pairs |
| Cost | Embeddings, storage and extra prompt tokens | Training runs, sometimes higher usage rates |
| Showing sources | Can cite the passages it used | Cannot point to where an answer came from |
| Getting started | Quick with managed tools | Slower: building the dataset takes most of the effort |
| Switching models | Works with any model | Tied to one base model |
How to choose between RAG and Fine-tuning
- Pick RAG when answers must come from your own content, stay current and show their sources.
- Pick fine-tuning when prompting cannot get a consistent style, format or narrow skill and you have good examples.
- Combine them when a tuned model should also answer from fresh documents.
The options
- RAGA way to make a language model answer from your own documents: search them for the passages relevant to a question, then hand those passages to the model with the question.
- Fine-tuningTraining an existing model further on your own examples so it picks up a particular style, format or narrow task, producing a customised version of that model.
More comparisons
- OpenAI vs Gemini vs Claude vs OpenRouterThree model makers and one gateway in front of them all. Prices are rough ranges per million input tokens; output tokens cost about four to six times more.
- Cloud API vs gateway vs running it yourselfThree ways to get a model's answers into an app: sign up with one provider, go through a gateway to many, or run an open model on hardware you control.
- Keyword vs vector vs hybrid searchKeyword search matches the words people type, vector search matches what they mean, and hybrid search runs both and merges the results.
- Node.js vs Deno vs BunThree runtimes for JavaScript and TypeScript on the server. Much of the same code runs on all three; they differ in built-in tools, security defaults, speed and how long each has been used in production.
- Express vs Fastify vs HonoThree JavaScript web frameworks with a similar feel. Express is the long-standing default, Fastify focuses on throughput and structure, and Hono is built on web standards so it can run almost anywhere.
- FastAPI vs Django vs FlaskThree widely used Python web frameworks. Django includes almost everything, Flask includes almost nothing, and FastAPI focuses on typed, self-documenting APIs.
Crafted in the dark. Shipped to the world.
Tell us what you are building. You get a private project space with a proposal and a line-by-line quote within a day.