Interesting MSR behavior: multiple characters inside one reference image can be individually controlled

#32
by ravibhai1212 - opened

Hi LiconStudio,

I have been testing LTX-2.3 MSR V2 in Wan2GP and discovered an interesting behavior that I could not find documented in the current examples.

Instead of using one reference image per character, I created a single reference image containing three different characters arranged spatially:

LEFT = Renu
CENTER = Kunal
RIGHT = Mukesh

I then used a separate image as the environment/background reference.

In the prompt, I explicitly defined the mapping:

Picture 1 = environment
Picture 2 = group reference
Left = Renu
Center = Kunal
Right = Mukesh

The result was surprisingly good.

All three characters appeared together in the generated video with their individual identities, clothing and general appearance preserved, even though they were all coming from the same composite reference image.

I also tested a two-character version, and that appears to be quite reliable.

The interesting part is the limitation I found:

When I tried to introduce the three characters sequentially over time:

  1. Renu appears first.
  2. Kunal enters from the side a few seconds later.
  3. Mukesh enters later.

Renu and Kunal remained fairly consistent, but Mukesh's facial identity drifted when he entered later in the sequence.

So it seems that a composite reference image can work as a kind of spatial identity map, but temporal introduction of a character may weaken identity fidelity.

I have attached screenshots/examples showing:

  • The original 3-character reference image
  • The environment reference
  • A successful video where all three characters appear together
  • The staggered-entry test where the third character's face changes

I think this could potentially be useful for improving MSR's identity anchoring.

It would be especially interesting if future versions could maintain the identity of each individual character from a composite reference image even when that character enters the scene later or is temporarily absent from the frame.

I would be happy to provide the exact prompts, reference images, model version, settings and generated videos if useful for reproduction.

Thank you for the amazing work on MSR!

001
002

2026-08-12_15h48_14

Hi LiconStudio,

I have been testing LTX-2.3 MSR V2 in Wan2GP and discovered an interesting behavior that I could not find documented in the current examples.

Instead of using one reference image per character, I created a single reference image containing three different characters arranged spatially:

LEFT = Renu
CENTER = Kunal
RIGHT = Mukesh

I then used a separate image as the environment/background reference.

In the prompt, I explicitly defined the mapping:

Picture 1 = environment
Picture 2 = group reference
Left = Renu
Center = Kunal
Right = Mukesh

The result was surprisingly good.

All three characters appeared together in the generated video with their individual identities, clothing and general appearance preserved, even though they were all coming from the same composite reference image.

I also tested a two-character version, and that appears to be quite reliable.

The interesting part is the limitation I found:

When I tried to introduce the three characters sequentially over time:

  1. Renu appears first.
  2. Kunal enters from the side a few seconds later.
  3. Mukesh enters later.

Renu and Kunal remained fairly consistent, but Mukesh's facial identity drifted when he entered later in the sequence.

So it seems that a composite reference image can work as a kind of spatial identity map, but temporal introduction of a character may weaken identity fidelity.

I have attached screenshots/examples showing:

  • The original 3-character reference image
  • The environment reference
  • A successful video where all three characters appear together
  • The staggered-entry test where the third character's face changes

I think this could potentially be useful for improving MSR's identity anchoring.

It would be especially interesting if future versions could maintain the identity of each individual character from a composite reference image even when that character enters the scene later or is temporarily absent from the frame.

I would be happy to provide the exact prompts, reference images, model version, settings and generated videos if useful for reproduction.

Thank you for the amazing work on MSR!

Thank you for the great testing feedback. I'll run some tests on this myself as well.

Hi LiconStudio,

I have been testing LTX-2.3 MSR V2 in Wan2GP and discovered an interesting behavior that I could not find documented in the current examples.

Instead of using one reference image per character, I created a single reference image containing three different characters arranged spatially:

LEFT = Renu
CENTER = Kunal
RIGHT = Mukesh

I then used a separate image as the environment/background reference.

In the prompt, I explicitly defined the mapping:

Picture 1 = environment
Picture 2 = group reference
Left = Renu
Center = Kunal
Right = Mukesh

The result was surprisingly good.

All three characters appeared together in the generated video with their individual identities, clothing and general appearance preserved, even though they were all coming from the same composite reference image.

I also tested a two-character version, and that appears to be quite reliable.

The interesting part is the limitation I found:

When I tried to introduce the three characters sequentially over time:

  1. Renu appears first.
  2. Kunal enters from the side a few seconds later.
  3. Mukesh enters later.

Renu and Kunal remained fairly consistent, but Mukesh's facial identity drifted when he entered later in the sequence.

So it seems that a composite reference image can work as a kind of spatial identity map, but temporal introduction of a character may weaken identity fidelity.

I have attached screenshots/examples showing:

  • The original 3-character reference image
  • The environment reference
  • A successful video where all three characters appear together
  • The staggered-entry test where the third character's face changes

I think this could potentially be useful for improving MSR's identity anchoring.

It would be especially interesting if future versions could maintain the identity of each individual character from a composite reference image even when that character enters the scene later or is temporarily absent from the frame.

I would be happy to provide the exact prompts, reference images, model version, settings and generated videos if useful for reproduction.

Thank you for the amazing work on MSR!

I’ve seen your test settings. I’ll test them as soon as possible.

Sign up or log in to comment