Back to Discover

#lora fine tuning

1 prompt found

MLX LM LoRA Fine Tuning Planner for Apple Silicon Macs: Quantized Base Model Choice, Chat JSONL Train and Valid Splits, mlx_lm.lora Flags for Memory, Prompt Masking, Test Loss, Adapter Fuse, and Before vs After Checks
๐Ÿค– AI Tools

MLX LM LoRA Fine Tuning Planner for Apple Silicon Macs: Quantized Base Model Choice, Chat JSONL Train and Valid Splits, mlx_lm.lora Flags for Memory, Prompt Masking, Test Loss, Adapter Fuse, and Before vs After Checks

PpromptstudioยทOct 7, 2026
No rating

Plan a LoRA fine tune of a small open model on your own Mac with mlx-lm: pick a quantized base that fits unified memory, convert your examples into chat JSONL train, valid, and test files, choose mlx_lm.lora flags for memory and loss masking, run evaluation on held out data, fuse the adapter, and compare outputs before and after on the same prompts.

Act as a machine learning engineer who fine tunes small open models on Apple Silicon with Apple's mlx-lm package, runs LoRA on 4 bit models with mlx_lm.lora, and has watched runs crash from memory pressure, learn the system prompt instead of the answer, and look good on training loss while failing held out examples. Inputs: - Mac model, chip, and unified memory: [MacSpecs] - The task in one sentence and what a correct output looks like: [TaskGoal] - Your examples: count, file type, and two real rows: [DatasetSample] - Base model you are considering, or let the plan pick one: [BaseModel] - How you will judge success (exact match on fields, a rubric, a reviewer): [EvalPlan] - Where the model will run afterwards (mlx_lm.generate, a local server, another runtime): [DeployTarget] - Output format: [Format] Generate: 1. A fit check: why the chosen base fits MacSpecs as a 4 bit MLX model from the mlx-community organization, and the smaller fallback if memory pressure appears. Training LoRA on a quantized base is the QLoRA setup. 2. A data plan: convert DatasetSample into chat format JSONL, one object per line with a messages list of system, user, and assistant turns, then split into data/train.jsonl, data/valid.jsonl, and data/test.jsonl with no duplicate rows across files. 3. A conversion script outline in Python that reads the source file, writes the three JSONL files, and prints counts and the longest example in tokens with the model tokenizer. 4. The mlx_lm.lora training command with each flag explained: --model, --train, --data, --batch-size, --num-layers, --iters, --learning-rate, --mask-prompt so loss is computed on the assistant answer only, --grad-checkpoint, --max-seq-length set from the longest example, --steps-per-eval, --val-batches, --save-every, and --adapter-path. 5. Memory moves in order if the run swaps or crashes: lower batch size, keep gradient checkpointing, cut num layers, shorten max sequence length, then switch to the smaller base. 6. Evaluation: the --test run on data/test.jsonl for test loss, plus a scripted pass that generates answers for every test row and scores them by EvalPlan. 7. A before versus after table on the same five test prompts with the base model and with the adapter. 8. Packaging for DeployTarget: mlx_lm.fuse with --save-path, and a note to run --help for current export options such as GGUF, which only some architectures support. Constraints: - Flags change between releases; tell the user to confirm with mlx_lm.lora --help. - Do not promise accuracy numbers or training times. No em dashes.