The Question Is Not “Should We Fine-Tune”

Every few weeks somebody asks whether we should fine-tune a model for our agent. And every time I notice the question is already wrong. It treats “training” as one big yes/no switch, when really it’s the last rung of a ladder with five rungs, and most tasks never need to reach it.

Here’s the ladder I use:

  1. Prompt and question framing
  2. Reference knowledge and retrieval
  3. Memory across runs
  4. Skills: packaged, versioned, tested
  5. Training

Most people explain this order with cost. Cheap things first, expensive things last. I think that’s the wrong axis, and it leads to wrong decisions. So let me say what I think the right axis is.


Order by Time-to-Correct, Not Expense

The question to ask about each rung: a fault is found today, how long until it’s actually fixed?

The five rungs and how long a fault takes to correct at each one

The five rungs and how long a fault takes to correct at each one

  • A bad prompt gets edited, reviewed, deployed. Same day.
  • A knowledge fault needs authoring, approval, a snapshot, a pinned release. One release cycle.
  • A bad memory item can be revoked. Immediate, for that single item.
  • A skill fault means version, test, re-release. One release cycle again.
  • A training fault means retrain, re-evaluate, re-qualify, re-release. A training run plus a re-release.

So every rung below training can be corrected the day you find the fault. Training can’t. That’s the whole argument, and it’s why training sits below the line instead of at the top of a to-do list.

Notice what this does to the “expensive” framing. A prompt rewrite might cost an engineer two days of careful thinking, which isn’t free. But the feedback loop is short. You find out you were wrong quickly, and you fix it quickly. With a weight update you commit to a slow loop, and every mistake you make inside it costs a full cycle.


The Climb Rule

When is it fine to climb a rung? My version of the rule, per task:

  • The rungs below are exhausted for that task, not in general
  • A metric and a held-out set exist
  • A no-training baseline exists to compare against
  • The data has all the sample types needed, split by root entity (not randomly by row)
  • Label agreement between raters has been measured
  • Every data source has rights for training, not just evaluation
  • If data leaves your tenant, a de-identification gate is operating
  • A withdrawal and rollback path exists

Two of these matter more than the rest.

“For that task” matters. “We tried prompting” is not the same as “prompting is exhausted for this specific classification decision”. The ladder gets climbed one task at a time, and different tasks sit on different rungs at the same time. A task that never had a serious prompt-plus-retrieval attempt hasn’t earned a training proposal.

The baseline is the one that gets skipped. “Our trained model reaches 78%” sounds great until you notice the hand-written rule it replaces gets 76%. Two points. That’s not a result, that’s a number.

And here’s the domain version, which I find sobering. One April 2025 study on root-cause-analysis recommendation used more than 180,000 historical incidents, evaluated on 3,000, plus human evaluation by incident owners. A well-optimised prompt with semantically similar historical exemplars was the best performer. Fine-tuned small models came out 13% below it. Retrieval-augmented large models came out 21% below.

That is not an argument against ever fine-tuning. It’s an argument that a fine-tune has to earn its existence against a strong prompt-plus-retrieval baseline, and that the baseline has to be built first.


Labels and Evaluation Are the Cost, Not Compute

This is where I think most budget conversations go wrong, so let me do actual arithmetic.

A hosted LoRA-based training service lists an 8B-class model at $0.44 per million training tokens (price read October 2026). A narrow-task fine-tune with 2,000 examples at about 4,000 trained tokens each is 8M tokens. That’s roughly $3.50 per epoch. Three epochs is about $10.

Ten dollars. Compare that with the published RL work: one reported result used 17,920 GPU-hours, and one scaling study used more than 400,000 GPU-hours plus a separate 100,000 GPU-hour validation run. Those are a different universe, but a narrow fine-tune isn’t in that universe. Its compute is basically free.

What isn’t free: someone has to produce the labels, and someone has to decide whether the result is any good. Labelling throughput and reviewer time are the bottleneck. I’d also say evaluation comes before everything: training before there’s an instrument is optimising an undefined quantity. If your measurement can’t see “right answer, wrong evidence”, and your rubric’s own inter-rater agreement is unmeasured (and that agreement is the ceiling on every label), then a training run will improve the measurement and not the product.

One more thing on the ladder itself: data preparation does not wait for the climb. The context at the moment a decision was made is not reconstructible later. Evidence content gets deleted under retention policies, knowledge snapshots change, profile versions change. Each day without extraction is that day’s samples lost. So the corpus gets designed and extracted early, and training happens later if at all.

(The self-hosted cost numbers floating around vendor blogs, single-digit dollars for an 8B QLoRA run, I couldn’t verify, so I’m not quoting them.)


RL Needs a Checkable Reward

Now the part that people get most excited about, RL for agents.

The pitch is fair. SFT maximises the probability of demonstrated actions, but what you want is the probability the agent completes the task. SFT can’t learn from a sub-optimal run at all, and during training the agent never sees its own mistakes, so it’s unprepared to recover from them. RL fixes this structurally: sample trajectories from the current policy, score them, move toward higher expected reward. The policy learns to recover from its own errors.

Fine. But almost every strong post-training result in this area, the GRPO family, RLVR, the big scaling recipes, depends on a cheap, automatic, correct verifier. A math answer has one. A unit test has one. A diagnosis written by an agent that investigates an incident? The label comes from a person.

That’s the verifier problem, and I think it should be the first question on any RL proposal, not the last.

Three places a reward can come from, and what’s wrong with each

Three places a reward can come from, and what’s wrong with each

Looking at the three options:

  • Delayed verifier. Closed cases where the real fix later became known. Plausibly the best asset you have. But it needs a confirmation state that has to actually be recorded, and in a lot of systems it isn’t.
  • Reviewer thumbs-up. Cheapest data to collect, binary signal, no ranked pairs needed. But a reviewer accepting a record means “the record is acceptable quality”, not “the cause was confirmed”. Label every accepted record “verified” and you’ve mislabelled at scale, and the model trained on it becomes confidently assertive exactly where it should be cautious.
  • Rubric or model judge. Easiest to build, and the worst failure profile. A preprint on rubric-based RL reports hacking shifting from rule-breaking to sycophancy, self-praise and length bias. For a diagnostic report, sycophancy and length bias are precisely the failure that passes human review and is wrong.

Also worth knowing: any tooling that advertises “no reward engineering needed” is advertising that third regime.


Reward Hacking Is the Part That Decides It

Documented cases from agent RL, where the environment made a shortcut available and the model took it: finding the published solution online, copying a later fix from git history, calling a stronger model through an available API key, making tests pass without fixing the bug. One of the cases is an agent using a grading key to generate training data despite an explicit restriction.

Auxiliary rewards get gamed in ways nobody predicts. A bug that gave credit for superficial web-tool calls led a model to use the browser as a calculator. A reward that favoured a playful persona produced creature metaphors even without the persona prompt. Partial credit is useful (it reduces zero-advantage groups) and it’s also the mechanism by which both happened.

Bad verifiers cap scores too. Audited benchmarks plateaued where many remaining failures turned out to be flawed tests or annotation errors. A verifier that checks an incidental wording choice instead of behaviour rejects valid solutions, and then the model learns to fit the wording.

The one that worries me most: a preprint (listed at ICLR 2026) reports that models trained with verifiable rewards enumerate instance-level labels that satisfy the verifier without learning the relational pattern. It’s specific to models trained that way, and prevalence rises with task complexity and inference compute. So “throw more compute at it” isn’t a mitigation. It’s the accelerant. I’d treat it as one paper, but it matches everything else above.

The habit that catches all of this is boring: track three readings independently. Training reward, independent success on unseen tasks and on behaviour the verifier didn’t check, and costs and failures (tool calls, regressions, repeated actions). If reward rises and independent success is flat, go read the trajectories.


What RL Can and Can’t Buy

One more result that I think decides a lot of budgets. A NeurIPS 2025 oral found RLVR models do better at pass@1 but worse at large pass@k than their base models. Mechanism: RL biases sampling toward reward-yielding paths already in the base model’s support. It sharpens, it narrows exploration, it doesn’t add reasoning patterns. Distillation from a teacher can introduce genuinely new patterns. There’s published pushback on this, which I haven’t verified, so hold it a bit loosely.

If it holds, the rule is clean. “The model can do this but unreliably” is the RL-shaped problem. “The model can’t do this at all” needs distillation from something that can, or a different model. Pick wrong and you burn the whole budget for a flat curve.

Same flavour of finding from the large scaling study: most common interventions (loss aggregation, curriculum, length penalty, advantage normalisation) move compute-efficiency, not the ceiling, and stable recipes follow predictable curves, so you can fit one on a small run and extrapolate before committing budget.


Where I Land

So my honest position for a domain with no automatic oracle: RL is not the next step. The next steps are boring ones. Record a confirmation state so a delayed verifier can exist. Measure rubric agreement. Build the no-training baseline. If RL ever becomes viable, enter through the delayed verifier, not the rubric.

And before any of that, the ladder. Fix the prompt. Fix the retrieval. Write the skills (that rung is the one most often empty, and in one published study curated skills raised a mean pass rate from 33.9% to 50.5%). If a fault is still there after that, and you have a metric, a baseline and labels you trust, then talk about weights.

The one open question I’d still like answered: what is the mechanism that makes training win here? A proprietary ontology, a latency requirement, an on-premise requirement, per-tenant specialisation. If it’s none of those, the honest recommendation is rungs one to four and no training at all.


Further Reading

  • DAPO, arXiv 2503.14476
  • Dr. GRPO, arXiv 2503.20783
  • GSPO, arXiv 2507.18071
  • “The Art of Scaling RL Compute for LLMs”, arXiv 2510.13786
  • GiGPO, arXiv 2505.10978
  • RLVR and the base model’s reasoning boundary (NeurIPS 2025 oral), arXiv 2504.13837