What is Fine-Tuning?
Definition
Fine-tuning is the process of continuing to train an already pre-trained AI model on a smaller dataset prepared for a specific task, style or domain. It changes the model's weights, which separates it from prompting, where only the input changes, and from RAG, where documents are retrieved at answer time. Typical goals are a consistent output format, a narrow classification task or a house writing style.
Also known as: model fine-tuning, fine tuning, supervised fine-tuning, SFT, LoRA

What actually changes
A foundation model picks up its general language ability in pre-training. Fine-tuning trains that model a little further on examples you supply, and its parameters shift to fit them. The result is a new version of the model that is more inclined to answer the way your examples do. The chat assistants people use every day are base models that their providers fine-tuned before release.
Three families of technique dominate:
- Supervised fine-tuning (SFT): learning from pairs of inputs and desired outputs. This is the most common route.
- Parameter-efficient methods: techniques such as LoRA train small add-on matrices instead of the full network, which cuts cost and hardware requirements considerably.
- Preference training: the model is shown pairs of answers to the same prompt, with one marked as better, and adjusts its tendencies accordingly.
A training example
Training data for chat models is usually stored as JSONL, one JSON object per line. Here is a single line from a dataset that teaches a model to route support tickets into a fixed set of queues:
{"messages": [
{"role": "system", "content": "Route the ticket to one queue: refunds, shipping, billing, other."},
{"role": "user", "content": "I was charged twice for order 4471."},
{"role": "assistant", "content": "billing"}
]}A few hundred to a few thousand clean examples like this can make the model reliably use the queue names and stick to one-word answers. A few dozen inconsistently labelled examples tend to do the opposite. Data quality matters far more than volume.
Prompting, RAG or fine-tuning?
| Approach | What changes | Good for | Weak spot |
|---|---|---|---|
| Prompting | Only the input | Quick iteration, varied tasks | Long instructions cost tokens on every call |
| RAG | Documents added to the context at answer time | Current facts that need citing | Poor retrieval means poor answers |
| Fine-tuning | The model's weights | Fixed formats, tone, narrow classification | A bad way to teach facts that change |
Most projects should start with a good prompt and a clear system prompt, add RAG when the task is knowledge-heavy, and consider fine-tuning only for behaviour problems that remain after both have been measured. The options combine well: a fine-tuned model can still read documents that RAG puts in front of it.
Where fine-tuning goes wrong
- Baking in volatile facts. Prices, stock levels or regulations written into the weights go stale the day they change. Keep that kind of information as a retrievable source, not as parametric knowledge.
- Expecting it to cure hallucination. Fine-tuning shapes style and behaviour; it does not on its own make the model factually reliable.
- Training without an evaluation. Without a held-out test set and a baseline (the prompt-only version) to compare against, there is no way to show the tuned model is better.
- Overfitting. Too many epochs on a small dataset can make the model parrot training examples and get worse at everything else.
There is also a maintenance cost. A fine-tune is tied to one base model version. When the provider retires that version, you have to rerun the training on its successor and test the results again.

