1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
When would you fine-tune a model for agentic behaviour instead of improving prompts and tools?
30-second answerSay your answer out loud first, then reveal.
Order of operations
- Better prompts, examples and tool descriptions.
- Better tools / fewer steps.
- Better context (retrieval, memory).
- A stronger base model.
- Then fine-tuning, if the economics and data justify it.
Good reasons to fine-tune
- Cost/latency: distil a large model's successful trajectories into a smaller model for a high-volume, narrow agent (e.g. ticket triage with 8 tools).
- Consistent domain behaviour: a specific tool-calling style, domain jargon, output formats that prompting can't hold reliably.
- Verifiable tasks + RL: where success can be checked automatically (tests pass, SQL returns the correct result), reinforcement fine-tuning can improve multi-step tool use.
Costs and risks
- Needs curated data (successful trajectories plus failures) and a robust eval suite.
- Lock-in: tool or schema changes may require retraining; new base models need re-tuning.
- Can lose general capabilities, or overfit to the training distribution.
- Ongoing MLOps: versioning, monitoring and retraining pipelines.
Interview framing: "I'd fine-tune only with an eval showing a clear gap that prompting and tooling couldn't close, plus enough volume to justify the operational cost."
Related
This is what real progress feels like.