# Why fine-tuning an agent is more than model tuning
# The four dials you must tune together Treat agentic fine-tuning as a system with four dials. Changing one without the others leaves failure modes.
- Training data: teach the behavior in the exact format the model will see at inference time. For tool-calling agents, format matters more than volume.
- Parameter-efficient fine-tuning: update model weights with techniques that fit available compute, like QLoRA and LoRA variants.
- Runtime hyperparameters: temperature, retry policy, and iteration limits are decided at inference and can break a well-trained agent if misconfigured.
- Preference alignment: judgment calls often require preference-based training (e.g., DPO) because single-label supervision cannot capture tradeoffs.
# Running example: a support-ticket triage agent The guide walks through a triage agent that must reliably call three internal tools: lookup_order, issue_refund, and escalate_to_human. The goal is that the agent emits syntactically exact tool calls with correct argument names rather than free-text answers.
# Parameter-efficient fine-tuning with QLoRA QLoRA is presented as a practical approach to fine-tune large models without needing datacenter-scale resources. The guide lists the local Python prerequisites for running examples: Python 3.10+, and packages such as peft, transformers, datasets, and accelerate. It notes that a real training run also needs bitsandbytes and a CUDA GPU where applicable.
After training, tune temperature, retry policy, and iteration limits with the same rigor as training hyperparameters. A well-trained model can still fail in production if inference-time temperature or retry logic encourages the wrong behavior.
# Teach judgment with Direct Preference Optimization (DPO) Supervised fine-tuning can teach exact formats and domain vocabulary but cannot always express tradeoffs and judgment calls. DPO is proposed to teach those preferences directly and make judgment-sensitive behavior consistent.
# Evaluate with a verdict-driven framework The article recommends a verdict-driven evaluation that catches catastrophic forgetting and other regressions. For the triage agent, that means verifying the model continues to call the correct tool with correct arguments across a representative test suite and flags any deviation as a failing verdict.
# Practical sequence
- 1Decide whether fine-tuning is the right tool. 2. Build and validate a small, well-structured tool-calling dataset. 3. Apply parameter-efficient fine-tuning (QLoRA/LoRA) given resource constraints. 4. Tune runtime hyperparameters empirically. 5. Use DPO for preference alignment where needed. 6. Run a verdict-driven evaluation before shipping.
This approach treats agentic fine-tuning as a holistic engineering task: all four dials must be tuned and validated together to produce reliable, production-ready agents.