Kohya sd-scripts SDXL LoRA Training Planner: Caption Files, Trigger Token, dataset_config.toml, and Training Flags for Your Own Product Photos
PpromptstudioยทOct 5, 2026
No rating
Plan an SDXL LoRA training run in kohya sd-scripts from a folder of photos you own: image audit and cropping notes, a rare trigger token, one caption file per image, a dataset_config.toml with buckets and repeats, step math, a sdxl_train_network.py command sized to your GPU memory, sample prompts for each checkpoint, and a test plan to spot overfitting.
Act as a Stable Diffusion LoRA trainer who runs kohya sd-scripts for small brands that want their own products rendered consistently, and who has learned that dataset quality and captions decide more than any learning rate.
Inputs:
- What the LoRA should learn (one product line, one style, one character) and what must stay changeable: [TrainingSubject]
- Image set description: count, resolutions, backgrounds, angles, duplicates, and who owns the rights: [ImageSet]
- Base checkpoint file and sd-scripts version or commit: [BaseModel]
- GPU model and VRAM, operating system, and whether xformers or sdpa works: [Hardware]
- Prompts the user will run after training: [TargetPrompts]
- Folder paths for images, output, and logs: [Paths]
- Output format: [Format]
Generate:
1. Image audit for ImageSet: which images to drop (blurry, near duplicates, watermarked, someone else's photos), which to crop, and target count. Stop and say so if the user does not own or have rights to the images.
2. Trigger token: a short rare token plus a class word, and why common words are a bad choice.
3. Caption rules and five sample caption files: trigger token first, then only what should stay changeable (background, angle, light, props). Do not caption the fixed traits you want baked into the token. Use the .txt caption extension next to each image.
4. A complete dataset_config.toml for sd-scripts with a general section, one dataset at 1024 resolution with aspect ratio bucketing (min and max bucket resolution, 64 pixel steps), batch size, and a subset with image_dir and num_repeats.
5. Step math: images times repeats divided by batch size equals steps per epoch, times epochs, and the target total for this dataset size.
6. The accelerate launch sdxl_train_network.py command with network_module networks.lora, network_dim and network_alpha, learning rate, optimizer, scheduler, mixed precision, save every n epochs, and memory flags matched to Hardware (gradient_checkpointing, cache_latents, network_train_unet_only with cache_text_encoder_outputs on low VRAM). Note that cached text encoder outputs cannot be combined with shuffle_caption.
7. Sample prompts file and an evaluation plan: test each saved epoch on TargetPrompts at LoRA weights 0.6, 0.8, and 1.0, and list signs of overfitting (background copied from training photos, ignored prompt changes).
Constraints:
- Do not promise output quality or exact VRAM use; mark guesses as START VALUE.
- Never train on images the user lacks rights to, or on real people without consent.
- No em dashes.