Prediction preview images with modular diffusers
I made some custom blocks as an example on how to achieve this with krea-2 and also published a space so you can watch it live.
What's the difference
Normally diffusers returns the latents of each denoise step, which are the image and the leftover noise mixed together. That's very useful to me in img2img: I can see exactly how much noise I'm adding, which makes it easier to adjust. The prediction can't show you that.
The other method is to return the prediction, which is what the model thinks the image is going to be at that specific step. It doesn't help me tune the noise, but I can tell early whether the image is going where I want and stop it if it isn't, instead of waiting for the whole thing. It also makes the waiting easier: there's a picture to watch from the first step instead of static.
The blocks
They're the normal Krea 2 blocks with one extra step added inside the denoise loop. That step builds the prediction and hands it to the callback function:
def on_preview(step, total, image):
image.save(f"preview_{step:02d}.png")
image = pipe(
prompt="a photograph of a cat wearing a wool hat",
preview_callback=on_preview,
output="images",
)[0]
What you get is a normal image, the same size as the final one, so you can save it or send it straight to a UI. If you don't pass the function, nothing happens and you get exactly the same result as always.
There are two ways to turn the prediction into that image. The default is a quick projection that costs nothing and looks blocky, good enough to tell where the generation is going. The other is a tiny decoder, 22 MB the first time you use it, that gives you real detail for about twenty milliseconds a step. That's the one the space uses as a default.
One nice addition is that you can stop a generation you don't like by raising an exception inside your function. It comes straight back out of the pipeline and you can generate again right after, which is what the stop button in the space does.
That's all there is to it. I hope you find it useful, and if you have an app built on diffusers I'd love to see this in it.
Credits
- The tiny decoder is TAEW 2.1 by Ollin Boer Bohan.
- The quick projection uses the latent-to-RGB factors from ComfyUI.
If you have any questions or suggestions, please don't hesitate to reach out to me on X OzzyGT or in the Diffusers channel on our Discord.

