Model customization

Fine-tuning

Continuing the training of an already trained model on a smaller set of your own examples, so that its weights shift towards a particular task, style or output format.

Base modelgeneral ability in place+ 3,000 examplespairs of input and ideal outputFine-tuned modelthe habit is set in the weightsBefore: large model, long prompt2,000 tokens · 94% right · $16,000 a monthAfter: small model, short prompt150 tokens · 95% right · $2,000 a monthIllustrative example. When the base model is retired, the tuning and the tests are run again.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/fine-tuning

In plain terms

A newly hired accountant already knows accounting. What they lack is your house style: how your reports are laid out, which phrases you use, what counts as an exception. A few weeks of worked examples settles that, and afterwards nobody has to explain it again. Fine-tuning does the same to a model. It learns a habit from examples and keeps it, with no reminder in the prompt. It is a poor way to teach facts: details taught this way are recalled unreliably.

Why it matters

Fine-tuning is often proposed first and should usually be tried last. It is the wrong tool when the gap is knowledge, since facts change and belong in documents the model reads; when there are only a few dozen examples, which fit in a prompt; and when the task is still being defined. It earns its cost when a stable, narrow job runs at high volume: a fixed format, a house tone, a classification, or a small model taught to do what a large one did. Budget for upkeep: when the base model is replaced or retired, the tuning and the testing are repeated.

Example

A logistics company turns 400,000 free-text delivery notes a month into a fixed twelve-field record. A large model with a 2,000-token prompt of rules and examples gets 94% right for 16,000 dollars a month. The team fine-tunes a small model on 3,000 corrected examples: 95% right, a 150-token prompt, 2,000 dollars a month. Eight months later the base model is retired, and the tuning and the tests are run again.

Most often confused with

Fine-tuning vs. RAG

Fine-tuningChanges how the model behaves; fixed until retrained
RAGChanges what the model reads; updated by editing a document

Ask where the gap is. If answers are wrong because the model has never seen your price list, training on the list gives patchy recall and a model that is out of date next month. If the facts are right but the tone, the structure or the judgement is off, more documents will not help. A fine-tuned model also cannot show where an answer came from. Mature systems often use both.

Under the hood

The common form is supervised fine-tuning: a file of input and ideal-output pairs, typically a few hundred to a few thousand, is run through the model for a small number of passes (epochs) at a low learning rate. Full fine-tuning updates every weight; parameter-efficient methods such as LoRA train a small add-on and leave the base untouched. Other forms: preference tuning on pairs of better and worse answers, reinforcement fine-tuning against a grader, and continued pre-training on unlabelled domain text. Hold back a test set and compare with the best prompt-only baseline. Pitfalls: overfitting, catastrophic forgetting of general ability, memorised training examples that can resurface (keep personal data and secrets out), and weakened safety behaviour. Availability for closed models varies and can change: in 2026 OpenAI told customers that its self-serve fine-tuning platform would stop accepting new training jobs in January 2027. Open-weight models can always be tuned on your own infrastructure.

Written by Mehmet Erkek · Last updated: