How to Use the MLX LoRA Fine Tuning Planner Prompt to Train a Small Model on Your Mac
Plan a LoRA fine tune of a small open model on your own Mac with mlx-lm: pick a quantized base that fits unified memory, convert your examples into chat JSONL train, valid, and test files, choose mlx_lm.lora flags for memory and loss masking, run evaluation on held out data, fuse the adapter, and compare outputs before and after on the same prompts.

Apple Silicon Macs can fine tune small open models with Apple's mlx-lm package, without renting a GPU. The common failures are predictable: the run crashes from memory pressure, the model learns to repeat the system prompt, or training loss looks great while held out examples still fail. The MLX LM LoRA Fine Tuning Planner for Apple Silicon Macs: Quantized Base Model Choice, Chat JSONL Train and Valid Splits, mlx_lm.lora Flags for Memory, Prompt Masking, Test Loss, Adapter Fuse, and Before vs After Checks prompt plans the whole job, from picking a base model that fits your memory to comparing outputs before and after the adapter.
What the prompt produces
- A fit check that picks a 4 bit base model from the mlx-community organization for your Mac's memory, with a smaller fallback.
- A data plan for chat format JSONL with train, valid, and test files and no duplicates across them.
- A conversion script outline that writes the files and reports the longest example in tokens.
- The mlx_lm.lora training command with each flag explained, including prompt masking, gradient checkpointing, and maximum sequence length.
- Memory moves in order if the run swaps or crashes.
- An evaluation plan using the test split for loss and a scripted pass scored by your own success rule.
- A before versus after table on the same test prompts.
- Packaging steps with mlx_lm.fuse and a note on export options.
How to fill the inputs
MacSpecs is the chip and unified memory, such as M2 Pro with 16 GB. Memory decides the base model size.
TaskGoal is one sentence on what the model should do and what a correct output looks like. Narrow tasks fine tune best.
DatasetSample gives your example count, file type, and two real rows. The conversion plan is built from these.
BaseModel can be a specific model or a size limit.
EvalPlan says how you will judge success. Exact field matches, a rubric, or a human reviewer all work if you state them.
DeployTarget is where the model runs next.
Reading the example output
The example teaches a model to turn plumbers' job notes into invoice line items as JSON on a 16 GB MacBook Pro:
- The base is sized to the machine: a 3B instruct model at 4 bit, with a 1.5B fallback if the Mac starts swapping.
- The split is explicit: 850 rows deduplicated, then 750 train, 50 valid, and 50 test with a fixed shuffle seed.
- Every flag has a reason. --mask-prompt keeps loss on the assistant JSON. --max-seq-length comes from the longest example the script prints, not a guess.
- The memory plan is ordered: batch size first, then fewer layers, then shorter sequences, then the smaller model.
- Evaluation goes beyond loss. A script parses each generated answer and checks four fields, and parse failures are counted separately.
- The before and after table is left blank for real results, with no promised accuracy.
Tips for better results
- Clean your examples before training. Inconsistent target formats teach inconsistent outputs.
- Watch validation loss at every eval step and keep earlier checkpoints.
- Run the base model on the test prompts first. Sometimes a better system prompt is enough.
- Confirm flags with mlx_lm.lora --help, since names change between releases.
Mistakes to avoid
- Do not let the same example appear in train and test.
- Do not set a long maximum sequence length just in case. It costs memory on every step.
- Do not judge success on training loss alone.
- Do not assume every export format is available for every model architecture.
Who it is for
Developers with Apple Silicon Macs who want a task specific model, small teams keeping data on their own hardware, students learning fine tuning hands on, and engineers testing whether a fine tune beats prompting before paying for cloud training.
Related PromptDig links
Open the MLX LM LoRA Fine Tuning Planner for Apple Silicon Macs: Quantized Base Model Choice, Chat JSONL Train and Valid Splits, mlx_lm.lora Flags for Memory, Prompt Masking, Test Loss, Adapter Fuse, and Before vs After Checks prompt and describe your task and data. For more AI tool and model prompts, Browse more prompts. If you have a fine tuning or local model prompt that works, Share a prompt.