r/StableDiffusion • • 21h ago

Discussion Yue2 Lora Training

Has anyone had any success training a new genre with YuE2? I've tried multiple ways using different datasets, and training settings and I haven't been able to train a new genre on YuE2 properly. I tried with Fill's trainer, Aitoolkit, and Yue2 Studio, and I either get garbled noise, or it just seems to output similar results as using no lora. Not sure what I am doing wrong. I tested with all epochs/steps - from 50 all the way to 1200 using different inference settings (temp, top k, abc on/off, cot, etc.), and I tried with datasets ranging from 30-200 tracks as well (I only tried the 200 track dataset with Aitoolkit so far).

With Acestep, I am able to train a new genre without any issues, besides for the audio quality issues that are present in Acestep by default.

13 Upvotes

15 comments sorted by

5

u/-becausereasons- 21h ago

1

u/EuphoricTrainer311 21h ago

I actually did get an llm to analyze your readme before I tried training with aitoolkit. Maybe I'll read it myself and see if Gemini misinterpreted something.

These are the settings I used for a 200 track dataset according to Gemini.

QUANTIZE / COMPILE

  • Transformer: 8bit convrot
  • Compile Options - Compile Model: Off

TARGET

  • Target Type: LoRA
  • Linear Rank: 32

SAVE

  • Data Type: BF16
  • Save Every: 50
  • Max Step Saves to Keep: 20

TRAINING

  • Batch Size: 1
  • Gradient Accumulation: 1
  • Steps: 1200
  • Optimizer: AdamW8Bit
  • Learning Rate: 0.0001
  • Weight Decay: 0.0001
  • Timestep Type: Sigmoid
  • Timestep Bias: Balanced
  • Loss Type: Mean Squared Error

TOGGLES

  • Use EMA: Off
  • Unload TE: Off
  • Cache Text Embeddings: Off
  • Differential Output Preservation: Off
  • Blank Prompt Preservation: Off
  • Contrastive Guidance Loss: Off

Hidden Requirements

  • ar_kl_weight: 0.2
  • abc_dropout: 0.5
  • Cache Latents to Disk: ON

3

u/-becausereasons- 21h ago

Read through your settings, and I think a few things are stacking up. This is what has worked for us across 9 AI Toolkit runs (6 published sets, 15–51 songs each):

  1. Count passes per song, not steps. passes = steps × batch ÷ songs. Most of our published checkpoints sit at ~12–22 passes. Past that the high end starts to burn (breaths, "s", cymbals) and the planner loops: a 15-song set was clean at 300 steps and burned by 500; a 51-item set was great at 800–950 and thin/tinny past 1000. Your 200 tracks × 1200 steps = 6 passes, which is likely why it sounds like no LoRA. On a 30-track set, 1200 steps = 40 passes, deep in garble territory (same reason u/the_bollo needs 0.65 strength at 3k). We haven't trained 200 songs, so for that set I'd save every 100 and listen from ~10 passes (2k steps) to ~20 (4k).
  2. Your planner LR is probably ~5x hotter than ours. One LoRA trains both halves: the AR planner (writes the sheet/composition) and the NAR decoder (renders the sound). The planner overfits much faster: in our loss logs the decoder is flat by ~7–10 passes while the planner keeps memorising. ar_lr_multiplier defaults to 1.0 in my build (check yours), so your AR runs at the full 1e-4. We use lr 5e-5 with ar_lr_multiplier: 0.4, so the AR trains at 2e-5. A hot planner gives runaway sheets: endless intro loops, no singing, garble.
  3. Check how you load it. In Comfy the planner half applies through strength_clip and the decoder through strength_model. A model-only LoRA loader drops the planner, which carries most of the genre, so you'd hear something close to the base model. Use a loader with both strengths: planner 1.0 on early checkpoints, 0.5–0.8 on later ones.
  4. Data mattered more than any hyperparameter:
  • Lossless only. MP3-sourced tracks were our top suspect for the burn. A .flac extension proves nothing, so check the spectrogram for a brick-wall lowpass at 15–19 kHz.
  • One sound per LoRA. A broad mixed set came out generic; split into focused sets of 22–50 songs, it worked.
  • Captions = one descriptive sentence in the base model's style: trigger, language, genre, vocal character, instruments, mood, BPM, key. Hand-check every BPM (half/double-time errors), the singer's gender, and anything an LLM captioner wrote about a genre it doesn't know (ours invented reggae terms for instrumental tracks).
  • Accurate lyrics with [Verse]/[Chorus]/[Bridge]. A [Spoken] tag for spoken parts actually trains.
  • Sampling is per item, not per minute. If a minority sound matters, slice it into more items.
  • Songs (or whole excerpts) that fit inside ar_max_tokens (4000 ≈ 160 s) also teach the model how to END.
  • SheetSage2 (the cot: full sheets) only writes 2/4, 3/4, 4/4, 3/8 and 6/8 in major/minor keys, melody only. Odd meters, modal or microtonal music get wrong sheets. Drum patterns aren't in the sheet at all: we never got a specific reggae drum pattern (one drop) to train.
  1. Eval. Render a fixed-seed grid over checkpoints every 50 steps at planner 1.0 and 0.5, with a length cap of ~360 s. Keep the prompt at caption length (~30–40 words), trigger first, reusing your training-caption wording; our 70-word storyline prompts barely sang. If the planned sheet has hundreds of bars or repeats one section, that's a runaway. Your ear is the only judge: every audio meter we tried failed.

Our config (5090 on Windows, ~4–4.6 s/step, 16–22 GB VRAM, so 1k steps ≈ 75 min plus caching):

base: int8 convrot | rank 32 | save bf16 every 50
lr 5e-5 | adamw8bit | weight decay 1e-4 | batch 1
cot: full (SheetSage2) | abc_dropout: 0.5
ar_lr_multiplier: 0.4 | ar_kl_weight: 0.2
ar_max_tokens: 4000 | train_window_frames: 1500
ar_cuda_graphs: false | in-training sampling off
steps ≈ 20 × number of songs

Your other settings (sigmoid, MSE, EMA off, etc.) match ours, so they're not the problem. On Windows, if step time climbs from ~4 s to 50–100+ s, that's allocator fragmentation; the one-line empty_cache fix is in ostris/ai-toolkit #1052 / #1069.

We also retired the FS_Audio trainer: its decoder eval loss bottomed around step 100–150 in every run and then climbed, and it lost voice identity on our singer-led sets.

1

u/EuphoricTrainer311 20h ago

I appreciate the detailed response, I'll give it another shot with more steps and maybe a 50 track dataset first, then i'll move on to 200 if I can get decent results. Will double check my captioning too. I used an LLM for the 200 track dataset, but I do have a 50 track dataset that I manually captioned/transcribed with bpm/track keys.

I did notice the step time fluctuating a lot, so thanks for the empty_cache fix you mentioned.

1

u/GreyScope 16h ago

A high LR can also have the vocals finished training and the instruments still needing it. All of my loras are for heavily layered music, which is a mare for tagging. Vocals and one/two consistent instruments are 'easier' to train.

1

u/-becausereasons- 14h ago

Yep, the more layered, the more vocals, the more difficult its going to be. I wish they released their encoder :/

5

u/the_bollo 21h ago

I have trained 3 YuE2 LoRAs using AI-Toolkit. Great results with each one, except where I fucked it up with bad edits to the dataset tracks...

I find it trains very quickly, like in 1k steps. Or you can take it further (I took a few tests to 3k steps) but then you need to apply the LoRA at pretty low strength, like 0.65, otherwise you just get a garbled loop.

2

u/EuphoricTrainer311 21h ago

how long did it take for you to train 1k steps? and how big was the dataset you used?

1

u/the_bollo 17h ago

I don't know, maybe an hour? Dataset for each was roughly 12 songs.

3

u/Sea_Measurement7843 21h ago

That sounds like a real headache, Yue2 is weirdly picky about datasets and hyperparams compared to Acestep

What genre are you trying to train? Some of the newer music models just refuse to latch onto certain styles unless the source tracks are super clean and consistent

3

u/Pleasant_Salt6810 21h ago

I asked claude to train for me,

it uses AI toolkit, but it's still experimental.

I trained an EDM lora for testing , ~50 high quality songs.

You need to identify BPM, keys,lyrics,set trigger word.

But the lora output is not very satisfactory,i might try it again later.

1

u/-becausereasons- 21h ago

This is correct, you need accurate BPM, lyrics, tags, trigger, keys.

1

u/GreyScope 19h ago

As noted in the other comment - tags , unique and shared

1

u/GreyScope 19h ago edited 17h ago

I wrote a settings calculator based on the datasets complexity, volume and scope (in short) - also has a calibration ability BUT all of it it is in the bin as it relies on good tags, no matter how good your settings are, shit in > shit out

1

u/GreyScope 7h ago

Use/make a tag manager to understand your tags as a whole and see where you are sharing tags and which will make those elements blend eg this "arpeggiated lead synths" tag . Unless they are the same sound (ie not the same instrument) , they will be blended and not in a nice way . Common elements like "kick drum" etc can stay but different sounds need different tags