🤖 AI Tools
MLX LM LoRA Fine Tuning Planner for Apple Silicon Macs: Quantized Base Model Choice, Chat JSONL Train and Valid Splits, mlx_lm.lora Flags for Memory, Prompt Masking, Test Loss, Adapter Fuse, and Before vs After Checks
Plan a LoRA fine tune of a small open model on your own Mac with mlx-lm: pick a quantized base that fits unified memory, convert your examples into chat JSONL train, valid, and test files, choose mlx_lm.lora flags for memory and loss masking, run evaluation on held out data, fuse the adapter, and compare outputs before and after on the same prompts.
0Reviews
Prompt
Act as a machine learning engineer who fine tunes small open models on Apple Silicon with Apple's mlx-lm package, runs LoRA on 4 bit models with mlx_lm.lora, and has watched runs crash from memory pressure, learn the system prompt instead of the answer, and look good on training loss while failing held out examples. Inputs: - Mac model, chip, and unified memory: [MacSpecs] - The task in one sentence and what a correct output looks like: [TaskGoal] - Your examples: count, file type, and two real rows: [DatasetSample] - Base model you are considering, or let the plan pick one: [BaseModel] - How you will judge success (exact match on fields, a rubric, a reviewer): [EvalPlan] - Where the model will run afterwards (mlx_lm.generate, a local server, another runtime): [DeployTarget] - Output format: [Format] Generate: 1. A fit check: why the chosen base fits MacSpecs as a 4 bit MLX model from the mlx-community organization, and the smaller fallback if memory pressure appears. Training LoRA on a quantized base is the QLoRA setup. 2. A data plan: convert DatasetSample into chat format JSONL, one object per line with a messages list of system, user, and assistant turns, then split into data/train.jsonl, data/valid.jsonl, and data/test.jsonl with no duplicate rows across files. 3. A conversion script outline in Python that reads the source file, writes the three JSONL files, and prints counts and the longest example in tokens with the model tokenizer. 4. The mlx_lm.lora training command with each flag explained: --model, --train, --data, --batch-size, --num-layers, --iters, --learning-rate, --mask-prompt so loss is computed on the assistant answer only, --grad-checkpoint, --max-seq-length set from the longest example, --steps-per-eval, --val-batches, --save-every, and --adapter-path. 5. Memory moves in order if the run swaps or crashes: lower batch size, keep gradient checkpointing, cut num layers, shorten max sequence length, then switch to the smaller base. 6. Evaluation: the --test run on data/test.jsonl for test loss, plus a scripted pass that generates answers for every test row and scores them by EvalPlan. 7. A before versus after table on the same five test prompts with the base model and with the adapter. 8. Packaging for DeployTarget: mlx_lm.fuse with --save-path, and a note to run --help for current export options such as GGUF, which only some architectures support. Constraints: - Flags change between releases; tell the user to confirm with mlx_lm.lora --help. - Do not promise accuracy numbers or training times. No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
MacSpecs: MacBook Pro 14 inch, M2 Pro, 16 GB unified memory TaskGoal: turn a plumber's free text job notes into invoice line items as JSON with description, quantity, unit, and labor hours, matching our office format exactly DatasetSample: 850 past jobs in jobs.csv with columns notes and invoice_json; example notes "replaced 40 gal water heater, new expansion tank, 2.5 hrs, haul away old unit" BaseModel: let the plan pick, under 4B parameters EvalPlan: exact match on each JSON field after parsing, plus office manager review of 20 random outputs DeployTarget: mlx_lm.generate on the office Mac first Format: steps, commands, script outline, eval table template
Example Output
1. Fit check
Base: mlx-community/Qwen2.5-3B-Instruct-4bit. A 3B model at 4 bit leaves room for LoRA training in 16 GB with small batches. Fallback if the Mac starts swapping: mlx-community/Qwen2.5-1.5B-Instruct-4bit. LoRA on this 4 bit base is the QLoRA setup.
2. Data plan
Each row becomes one line:
{"messages": [{"role": "system", "content": "Convert job notes to invoice line items as JSON."}, {"role": "user", "content": "<notes>"}, {"role": "assistant", "content": "<invoice_json>"}]}
Split 850 rows after a fixed shuffle: 750 train, 50 valid, 50 test. Deduplicate on notes before splitting so the same job never lands in two files.
3. Conversion script outline (prepare_data.py)
- Read jobs.csv with the csv module.
- Parse invoice_json; skip and log rows that do not parse.
- Drop duplicate notes, shuffle with random.seed(7), slice 750 / 50 / 50.
- Write data/train.jsonl, data/valid.jsonl, data/test.jsonl with json.dumps per line.
- Load the tokenizer with mlx_lm.load, apply the chat template to each row, print counts and the longest example in tokens.
4. Training command
pip install -U mlx-lm
mlx_lm.lora \
--model mlx-community/Qwen2.5-3B-Instruct-4bit \
--train --data ./data \
--batch-size 2 --num-layers 16 --iters 600 \
--learning-rate 1e-5 --mask-prompt --grad-checkpoint \
--max-seq-length 1024 \
--steps-per-eval 100 --val-batches 25 --save-every 100 \
--adapter-path ./adapters
- --mask-prompt: loss only on the assistant JSON, so the model does not spend capacity learning to repeat the system prompt and notes.
- --max-seq-length 1024: set from the longest example printed in step 3, rounded up. Raise it only if that number is higher.
- --num-layers 16: applies LoRA to the last 16 layers; fewer layers use less memory.
- --save-every 100: checkpoints so a crash does not lose the run.
- Watch validation loss at each eval. If it rises while training loss keeps falling, stop and use an earlier checkpoint.
5. If memory pressure appears
1. --batch-size 1
2. keep --grad-checkpoint
3. --num-layers 8
4. --max-seq-length 768 if the longest example allows
5. switch to the 1.5B base
6. Evaluation
mlx_lm.lora --model mlx-community/Qwen2.5-3B-Instruct-4bit --adapter-path ./adapters --data ./data --test
Then eval.py: for each test row, generate with the adapter, json.loads the output, and compare description, quantity, unit, and labor_hours to the reference. Count parse failures separately. Send 20 random outputs to the office manager with the original notes.
7. Before vs after (fill from your runs)
| Test prompt | Base output valid JSON? | Base fields matched | Adapter valid JSON? | Adapter fields matched |
| water heater job | [ ] | [ /4 ] | [ ] | [ /4 ] |
| drain clearing | [ ] | [ /4 ] | [ ] | [ /4 ] |
| fixture swap | [ ] | [ /4 ] | [ ] | [ /4 ] |
| leak repair | [ ] | [ /4 ] | [ ] | [ /4 ] |
| emergency call | [ ] | [ /4 ] | [ ] | [ /4 ] |
8. Packaging
mlx_lm.generate --model mlx-community/Qwen2.5-3B-Instruct-4bit --adapter-path ./adapters --prompt "<notes>" --max-tokens 300
To ship one folder: mlx_lm.fuse --model mlx-community/Qwen2.5-3B-Instruct-4bit --adapter-path ./adapters --save-path ./fused-invoice-model
Run mlx_lm.fuse --help before planning a GGUF export; it is only offered for some architectures. Confirm every flag above with mlx_lm.lora --help for your installed version.