🤖 AI Tools
Axolotl QLoRA Fine Tuning Config Builder: Chat Template Dataset Mapping, Assistant Only Loss, Sample Packing, Batch Math for One GPU, Preprocess Checks, and Merge Commands
Build an Axolotl YAML config for a QLoRA fine tune of an open instruct model on one GPU: a dataset format check for chat_template messages, assistant only training, LoRA and quantization settings, sequence length and sample packing, effective batch and step math, out of memory fallbacks, the preprocess, train, inference, and merge commands, and an evaluation plan against the base model.
0Reviews
Prompt
Act as an ML engineer who fine tunes open weight models with Axolotl for small teams, and who writes configs that run the first time on the hardware the team actually has. Inputs: - Base model from the Hugging Face Hub and its license terms you have accepted: [BaseModel] - GPU model, count, and VRAM: [GPU] - The task the tuned model should do and the exact output it should produce: [TaskGoal] - Two or three real dataset rows as they appear in the file: [DatasetSample] - Row count and typical length of a conversation: [DatasetSize] - How success will be judged: [EvalPlan] - Output format: [Format] Generate: 1. A dataset check: confirm DatasetSample is JSONL with a messages list of role and content turns. Flag rows with missing roles, empty assistant turns, or outputs that do not match TaskGoal, and give a one line fix for each problem. 2. A complete Axolotl YAML: base_model, load_in_4bit with adapter: qlora, lora_r, lora_alpha, lora_dropout, lora_target_linear, a datasets entry with type: chat_template, field_messages, roles_to_train set to assistant and train_on_eos: turn, dataset_prepared_path, val_set_size, output_dir, sequence_len chosen from DatasetSize, sample_packing, micro_batch_size, gradient_accumulation_steps, num_epochs, learning_rate, optimizer, lr_scheduler, warmup, bf16, gradient_checkpointing, flash_attention, evals_per_epoch, saves_per_epoch. 3. Batch math: effective batch = micro_batch_size x gradient_accumulation_steps x GPU count, and steps per epoch from the training rows before packing. Note that packing lowers the step count and that Axolotl reports the real number after preprocessing. 4. Out of memory fallbacks for GPU in order: smaller micro_batch_size with more accumulation, shorter sequence_len, then a smaller base model. 5. Commands: axolotl preprocess with --debug to confirm only assistant tokens carry labels, axolotl train, axolotl inference with --lora-model-dir, and axolotl merge-lora. 6. An evaluation plan from EvalPlan: a held out test file that never appears in training, the same system prompt for base and tuned model, and the metric for each output field. Constraints: - Do not invent accuracy gains, loss values, or training times. Say what to measure instead. - Keep personal data out of training rows, or say it must be removed first. - Check key names against the Axolotl version installed, since options are renamed between releases. No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
BaseModel: Qwen/Qwen2.5-7B-Instruct
GPU: one RTX 4090, 24 GB
TaskGoal: read a customer support ticket and return JSON with queue (billing, shipping, technical, account) and priority (P1 to P3), nothing else
DatasetSample: {"messages": [{"role": "system", "content": "Route the ticket. Reply with JSON only."}, {"role": "user", "content": "I was charged twice for my October order."}, {"role": "assistant", "content": "{\"queue\": \"billing\", \"priority\": \"P2\"}"}]}
DatasetSize: 6,000 rows, most tickets under 400 tokens
EvalPlan: 300 held out tickets labeled by the support lead; measure valid JSON rate, queue accuracy, priority accuracy
Format: checks, YAML, math, commands, evalExample Output
Dataset check
The sample row is valid: system, user, and assistant turns, and the assistant content is JSON only. Before training, drop rows where the assistant text has anything besides the JSON object, and remove names, emails, and order numbers from ticket text.
ticket-triage.yml
base_model: Qwen/Qwen2.5-7B-Instruct
load_in_4bit: true
adapter: qlora
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_linear: true
chat_template: tokenizer_default
datasets:
- path: data/tickets_train.jsonl
ds_type: json
type: chat_template
field_messages: messages
roles_to_train: ["assistant"]
train_on_eos: turn
dataset_prepared_path: last_run_prepared
val_set_size: 0.05
output_dir: ./outputs/ticket-triage-qlora
sequence_len: 1024
sample_packing: true
pad_to_sequence_len: true
micro_batch_size: 2
gradient_accumulation_steps: 8
num_epochs: 2
learning_rate: 0.0002
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
warmup_ratio: 0.05
bf16: auto
gradient_checkpointing: true
flash_attention: true
logging_steps: 10
evals_per_epoch: 4
saves_per_epoch: 1
Why 1024: tickets are under 400 tokens, so 1024 leaves room for the system prompt and packing fits two or more tickets per sequence.
Batch math
Training rows: 6,000 minus 5 percent validation = 5,700.
Effective batch: 2 x 8 x 1 GPU = 16.
Steps per epoch before packing: 5,700 / 16 = about 357, so about 713 for 2 epochs. Packing lowers this; use the count Axolotl prints after preprocessing.
If you run out of memory
1. micro_batch_size: 1 and gradient_accumulation_steps: 16 (same effective batch).
2. sequence_len: 768.
3. Switch to a 3B instruct model with the same config.
Commands
axolotl preprocess ticket-triage.yml --debug
Check the debug view: system and user tokens should be masked (label -100) and only the assistant JSON and its end of turn token trained.
axolotl train ticket-triage.yml
axolotl inference ticket-triage.yml --lora-model-dir="./outputs/ticket-triage-qlora"
axolotl merge-lora ticket-triage.yml --lora-model-dir="./outputs/ticket-triage-qlora"
Evaluation
Keep the 300 labeled tickets in data/tickets_test.jsonl and confirm none appear in the training file. Run the base model and the tuned model with the same system prompt. Report valid JSON rate, queue accuracy, and priority accuracy for each, plus a confusion table for queue so you can see which queues get mixed up.