2 comments

  • schopra909 24 minutes ago
    The approach is cool; but the results feel ~sus~.

    GenEval and GenEval2 prompts are very terse (~10 tokens) while these models are primarily trained on long-dense captions. So, it's pretty common to upsample prompts before running them through benchmarks. That way you're testing the model as it was trained, rather than testing it on prompts that are out of distribution.

    If you look at the Qwen-Image Report before they enhanced it for their 12-25 release, upsampled prompts score 0.87 on GenEval* In this paper they take it from 0.74 to 0.81.

    To me that reads to me that they're basically getting the model to adapt to the benchmarks' terse prompts. And it's unclear to me if simply finetuning the model on shorter prompts would work just as well as the RL solution.

    * https://arxiv.org/pdf/2508.02324v1#page=21

  • Lerc 1 hour ago
    Interesting, I had been wondering if you could cycle distillation and then splitting weights to turn a saturated model into a non-saturated with the same parameter count for further training.

    Train until you stop getting sidnificant improvements. distill to a quarter size model. Expand back up, and continue training.

    Split the weights so W1+W2 = W with W1 = (W + Randoffset)/2, W2 = (W - RandOffset)/2. Double the width of the layers using the split weights, you get the same result from a 4x size network. My hypothesis is that this has way more scope to train that the model you distilled from.