myibrahim.cloud

AI / ML · Tuning & Eval

When fine-tuning beats prompt engineering

Most fine-tuning attempts are premature. Here's how to tell when you actually need it, and what cheaper alternatives to try first.

The instinct, when an LLM behaves badly on your task, is to fine-tune. We'll teach it our domain. Most of the time, this is wrong. Fine-tuning costs more than the engineers expect, takes longer than the calendar shows, and rarely gives the lift the original prompt-engineering session skipped.

This is the framework I use to decide.

The cheaper alternatives, in order#

Before fine-tuning, you should have tried — and measured — all of these:

  1. A more capable model. Sometimes Claude Opus does what Sonnet can't. The price-per-token is higher; the engineer-hour-per-feature is lower.
  2. Better prompts. Few-shot examples. Explicit reasoning steps. XML-tagged structured input. Most "the model can't do X" turn out to be "the prompt didn't ask for X clearly."
  3. Retrieval (RAG). If the model needs domain knowledge it doesn't have — your product docs, your internal vocabulary — fine-tuning is rarely the right tool. RAG is.
  4. Tool use. If the model needs to do something — query a DB, call an API — give it a tool. Don't try to make it memorize the data.
  5. Prompt caching. Same prompt, 90% cheaper, no fine-tuning required.

Only after all five have been measured-and-found-wanting does fine-tuning earn its place on the table.

When fine-tuning genuinely helps#

There are real cases. They are narrower than people assume.

Style and format consistency. If you need outputs in a very specific style — legal contracts in a particular firm's voice, summaries with a precise structure, code in a house-specific framework — and prompts can't get you there reliably across thousands of calls, fine-tuning anchors style.

Latency on cheap hardware. A fine-tuned 7B model running on your own GPU, returning in 50ms, can replace a 200B-param API call for a narrow task. The cost crossover happens when you have ~100K calls/day at the same quality.

Refusing or accepting differently. A safety-focused fine-tune that refuses certain inputs your generic API model won't, or accepts certain inputs the generic model overzealously refuses.

Following a domain-specific instruction set. Tax filing assistants. Medical coding assistants. Where the task itself is highly structured and a generic model needs lots of pre-amble to perform.

In every case the question to ask first is: can I show in an eval that prompts max out before the model does? If yes, fine-tune. If no, fine-tuning won't fix it either.

The two-axis pricing matrix#

Approx. fine-tuning cost ($) vs. examples (Claude tier)
  • 1k2
  • 10k18
  • 100k180
  • 1M1,800

Fine-tuning costs scale with two things: how big the dataset is, and how many epochs you train. Across providers, ballpark numbers as of mid-2026:

Approach Setup cost Per-call premium Quality lift
OpenAI/Anthropic hosted fine-tune (1K–10K examples) $50–500 2–4× base price small to medium
LoRA/QLoRA on open model (10K examples) $20 (1 hour A100) self-hosted medium, narrow
Full fine-tune of 7B open model $100–500 self-hosted medium-high
Full fine-tune of 70B open model $5K–20K self-hosted high

The hosted-fine-tune option is great when "I want a slightly better-formatted version of this model." The self-host option starts paying off at high volume.

What good fine-tuning data looks like#

The mistake most teams make: dumping their entire backlog into a JSONL file and submitting. Models trained on noisy data become noisy models.

A good fine-tuning dataset is:

  • Small — 200 high-quality examples beat 10K low-quality. Yes, really.
  • Diverse — covers the input distribution you actually see in production, including edge cases.
  • Label-clean — every output is something you'd be proud to ship. Low quality outputs poison the model.
  • Format-consistent — if you want the model to always emit JSON, every training example is JSON. No exceptions.
  • Eval-segregated — at least 20% held out for evaluation, never seen in training.

A pragmatic loop: log production traffic, hand-curate the best 200–500 (input, output) pairs, train, eval, iterate.

LoRA and QLoRA: cheap fine-tuning that works#

LoRA (Low-Rank Adaptation) freezes the original model weights and trains a small "adapter" matrix on top. Models trained with LoRA can be ~1000× smaller than full fine-tunes while capturing 80–95% of the quality.

QLoRA quantizes the base model to 4-bit during training, fitting a 70B model on a single 80GB A100. Costs drop from "team's quarterly budget" to "weekend project."

# QLoRA on Llama-3-70B with the trl library
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig
from trl import SFTTrainer

bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype="float16")
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3-70B-Instruct",
    quantization_config=bnb,
    device_map="auto",
)

lora = LoraConfig(r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"])
trainer = SFTTrainer(model, args=training_args, peft_config=lora, train_dataset=ds)
trainer.train()

For most teams this is the right entry point.

How to tell if it worked#

Run the same eval set against the base model and the fine-tuned model. Three things to measure:

  1. Task accuracy on held-out data. Fine-tune should be ≥ base model.
  2. General capability regression. Run a generic benchmark (MMLU subset, HumanEval) — fine-tune shouldn't be much worse.
  3. Hallucination rate on out-of-distribution inputs. Fine-tunes often hallucinate more confidently. Test with junk inputs and check for sensible "I don't know" responses.

If accuracy went up but capability regressed badly, your training data is too narrow. Mix in some general examples or use a smaller learning rate.

The hardest part is the data#

Fine-tuning works when you can produce a clean, diverse, well-labeled dataset. The actual training step is mostly a trainer.train() call. The hard part — the part that takes weeks — is the data.

If your team can't produce 500 hand-curated examples, you're not ready to fine-tune. Spend that week improving the prompt instead.

Further reading#

  • ai
  • fine-tuning
  • llms
  • training
  • evals
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.