Using Krea with reference images

Community Article
Published September 29, 2026

Krea-2 has a really good feature, which is that it uses the Qwen3-VL encoder. This allows us to pass an image along with the prompt. In simple terms, we can just skip the whole description of an image we want and pass one instead.

Currently, I know of three ways of doing this, each one for a different use case:

  • Vision only
  • Style Reference with LoRA
  • Identity Edit with LoRA

This is a brief post about the three modes we can use with Krea-2, using Modular Diffusers, custom blocks, LoRAs, and SDNQ. If you just want to play with them, you can use this Space, or if you just want the code, you can go to the Diffusers recipes and get it there.

Vision only

This method doesn't require any LoRA. We just need to encode the image and pass the embeddings to the model along with the prompt.

Using this image as a reference:

Reference Image

Original photo by Microsoft Copilot on Unsplash

With just this code:

import torch
from sdnq import SDNQConfig  # noqa: F401

from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image


blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.to("cuda")

image = pipe(
    prompt="photo",
    reference_images=load_image(
        "https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/20260907154202_edit_image_0.png"
    ),
    output="images",
)[0]

which basically uses the the prompt "photo" and the image, we get this result:

Result Image

As you can see, it does a pretty good job, but it also misses some details that the model probably can't see. So if you want a more precise replica of the image, you can add the missing details to the prompt, like ethnicity or what each element on the bed is. Using this method, you can also do some very light editing of the image.

For example, with this prompt: "change the color of the sweater to light blue and replace the popcorn with french fries and scones with a hamburger" you will get this image:

Result Image

As you can see, it does change the image somehow, but every time we do it, the model produces a different result, and it's also not as precise. This is equivalent to prompting something like this: "a photo of an orange car in the city, change the color of the car to red", so it does kind of make sense that it works, with the first part of the prompt being the embeddings from the image.

One cool feature that I haven't seen anyone using is that, if you use a mask, you can prevent the model from seeing the parts that you don't want or need. So, for example, if I mask everything but the man, we get this image:

Result Image

Or if I mask just the laptop:

Result Image

The script used for these images is this one: reference_vision_only.py

Style reference

One thing that doesn't work with just vision is transferring the style from the image to the generation. For this, Ostris trained a LoRA that does this pretty well.

This one is pretty straightforward to use. You just prompt like normal and it will use the style of the image.

Using this image:

Reference Image

Photo by HYEWON HWANG on Unsplash

with this code:

import torch
from sdnq import SDNQConfig  # noqa: F401

from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image


blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.load_lora_weights("ostris/krea2_turbo_style_reference", weight_name="krea2_style_reference.safetensors")
pipe.to("cuda")

image = pipe(
    prompt="a capybara",
    reference_images=load_image(
        "https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/20260908045229_edit_image_0.png"
    ),
    reference_mode="append",
    output="images",
)[0]

Will produce this image:

Result Image

It kind of works, but this is a bad image to test it with since it has three subjects too close to each other, so it mixes them into one. This is where it gets interesting to use a mask, since you can isolate the specific style and subject you want from any image. If we mask only the red plushie in the middle, we get this image:

Result Image

Pretty cool, and it allows for some really creative experimentation, since you don't need to mask a specific subject or a clearly defined part of an image.

The script used for these images is this one: reference_style.py

Identity edit

This one is just like any other edit model. You pass an image and ask the model to edit something in it, and it should preserve most of the image.

Using this image:

Reference Image

Photo by Alvin David on Unsplash

and this code:

import torch
from sdnq import SDNQConfig  # noqa: F401

from diffusers.modular_pipelines import ModularPipelineBlocks
from diffusers.utils import load_image


blocks = ModularPipelineBlocks.from_pretrained("OzzyGT/krea2_reference_blocks", trust_remote_code=True)
pipe = blocks.init_pipeline("OzzyGT/Krea_2_Turbo_sdnq_dynamic_8bit")
pipe.load_components(dtype=torch.bfloat16)
pipe.load_lora_weights("conradlocke/krea2-identity-edit", weight_name="krea2_identity_edit_v1_2.safetensors")
pipe.to("cuda")

image = pipe(
    prompt="add sunglasses and a hat to the capybara",
    reference_images=load_image(
        "https://huggingface.co/datasets/OzzyGT/diffusers-examples/resolve/main/krea2_reference/alvin-david--cHj_FC6ukw-unsplash_croped.jpg"
    ),
    reference_mode="prepend",
    mask_reference_latents=True,
    output="images",
)[0]

we get this image:

Result Image

I don't see that much use for masks with this one, but you still can if you want, for example, to crop the capybara out of the image.

Result Image

The script used for these images is this one: reference_identity_edit.py

Hope you liked it! These three techniques can help improve your workflow and add a little extra to your generations.

If you have any questions or suggestions, please don't hesitate to reach out to me on X OzzyGT or in the Diffusers channel on our Discord.

Community

krea is really good.

Sign up or log in to comment