🤖 AI Tools

Kohya sd-scripts SDXL LoRA Training Planner: Caption Files, Trigger Token, dataset_config.toml, and Training Flags for Your Own Product Photos

Plan an SDXL LoRA training run in kohya sd-scripts from a folder of photos you own: image audit and cropping notes, a rare trigger token, one caption file per image, a dataset_config.toml with buckets and repeats, step math, a sdxl_train_network.py command sized to your GPU memory, sample prompts for each checkpoint, and a test plan to spot overfitting.

0.0
0Reviews
P
October 5, 2026

Prompt

Act as a Stable Diffusion LoRA trainer who runs kohya sd-scripts for small brands that want their own products rendered consistently, and who has learned that dataset quality and captions decide more than any learning rate.

Inputs:
- What the LoRA should learn (one product line, one style, one character) and what must stay changeable: [TrainingSubject]
- Image set description: count, resolutions, backgrounds, angles, duplicates, and who owns the rights: [ImageSet]
- Base checkpoint file and sd-scripts version or commit: [BaseModel]
- GPU model and VRAM, operating system, and whether xformers or sdpa works: [Hardware]
- Prompts the user will run after training: [TargetPrompts]
- Folder paths for images, output, and logs: [Paths]
- Output format: [Format]

Generate:
1. Image audit for ImageSet: which images to drop (blurry, near duplicates, watermarked, someone else's photos), which to crop, and target count. Stop and say so if the user does not own or have rights to the images.
2. Trigger token: a short rare token plus a class word, and why common words are a bad choice.
3. Caption rules and five sample caption files: trigger token first, then only what should stay changeable (background, angle, light, props). Do not caption the fixed traits you want baked into the token. Use the .txt caption extension next to each image.
4. A complete dataset_config.toml for sd-scripts with a general section, one dataset at 1024 resolution with aspect ratio bucketing (min and max bucket resolution, 64 pixel steps), batch size, and a subset with image_dir and num_repeats.
5. Step math: images times repeats divided by batch size equals steps per epoch, times epochs, and the target total for this dataset size.
6. The accelerate launch sdxl_train_network.py command with network_module networks.lora, network_dim and network_alpha, learning rate, optimizer, scheduler, mixed precision, save every n epochs, and memory flags matched to Hardware (gradient_checkpointing, cache_latents, network_train_unet_only with cache_text_encoder_outputs on low VRAM). Note that cached text encoder outputs cannot be combined with shuffle_caption.
7. Sample prompts file and an evaluation plan: test each saved epoch on TargetPrompts at LoRA weights 0.6, 0.8, and 1.0, and list signs of overfitting (background copied from training photos, ignored prompt changes).

Constraints:
- Do not promise output quality or exact VRAM use; mark guesses as START VALUE.
- Never train on images the user lacks rights to, or on real people without consent.
- No em dashes.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Kohya sd-scripts SDXL LoRA Training Planner: Caption Files, Trigger Token, dataset_config.toml, and Training Flags for Your Own Product Photos - Result

Examples

Example Input

TrainingSubject: our handmade speckled stoneware mugs (shape, speckle pattern, matte white glaze) so we can render them in new scenes; background, props, and light must stay changeable
ImageSet: 34 photos taken by our studio, 3000x4000 phone shots, 20 on a white sweep, 14 lifestyle on wood tables, 4 near duplicates, 2 slightly blurry
BaseModel: sd_xl_base_1.0.safetensors, sd-scripts recent main branch
Hardware: 12 GB VRAM, Linux, sdpa works
TargetPrompts: mug on a cafe counter in morning light; mug held in two hands by a window; flat lay with pastries
Paths: /data/fwmug/img, /data/fwmug/out, /data/fwmug/logs

Example Output

1. Image audit
Drop 4 near duplicates and 2 blurry shots: 28 images remain. Crop the white sweep shots so the mug fills about 60% of the frame. Rights: your studio photos, OK.

2. Trigger token: fwmug mug
A rare token like fwmug starts with little meaning in the model; a common word like speckle would fight what the model already knows.

3. Sample captions (one .txt per image)
IMG_0412.txt: fwmug mug, on a white background, front view, soft studio light
IMG_0418.txt: fwmug mug, on a white background, three quarter view, handle on the right
IMG_0533.txt: fwmug mug, on a wooden table, coffee inside, window light, plant in background
IMG_0540.txt: fwmug mug, on a wooden table, overhead view, linen napkin
IMG_0551.txt: fwmug mug, held in hands, warm evening light
Not captioned on purpose: speckles, matte white glaze, shape. Those belong to fwmug.

4. dataset_config.toml
[general]
caption_extension = ".txt"
shuffle_caption = false

[[datasets]]
resolution = 1024
batch_size = 2
enable_bucket = true
min_bucket_reso = 640
max_bucket_reso = 1536
bucket_reso_steps = 64

  [[datasets.subsets]]
  image_dir = "/data/fwmug/img"
  num_repeats = 10

5. Step math
28 images x 10 repeats / batch 2 = 140 steps per epoch. 10 epochs = 1,400 steps (START VALUE; a single product often settles between 1,000 and 2,000).

6. Command
accelerate launch sdxl_train_network.py \
  --pretrained_model_name_or_path=/models/sd_xl_base_1.0.safetensors \
  --dataset_config=dataset_config.toml \
  --output_dir=/data/fwmug/out --output_name=fwmug_v1 --logging_dir=/data/fwmug/logs \
  --network_module=networks.lora --network_dim=32 --network_alpha=16 \
  --learning_rate=1e-4 --optimizer_type=AdamW8bit --lr_scheduler=cosine \
  --max_train_epochs=10 --save_every_n_epochs=2 --save_model_as=safetensors \
  --mixed_precision=bf16 --save_precision=bf16 --sdpa \
  --gradient_checkpointing --cache_latents \
  --network_train_unet_only --cache_text_encoder_outputs \
  --sample_prompts=sample_prompts.txt --sample_every_n_epochs=2 --seed=42
shuffle_caption is false because cached text encoder outputs cannot be shuffled.

7. sample_prompts.txt
fwmug mug on a cafe counter, morning light --w 1024 --h 1024 --s 28
fwmug mug held in two hands by a window --w 896 --h 1152 --s 28
Evaluation: load epochs 4, 6, 8, 10 at weights 0.6, 0.8, 1.0 on the three TargetPrompts. Pick the earliest epoch where speckles and shape hold. Overfitting signs: the wooden table appears when you asked for a counter, or the white sweep leaks into every scene.

Reviews (0)

Please login to leave a review.
Loading reviews...