← Back to OmniFit Blog
Hands-on workflow notes · MiniMax H3 Ref2VA

What worked, what broke, and which prompt template to start from

📅 Published: August 16, 2026

These notes come from real local Ref2VA tests, not a one-shot demo. We tried single-person swaps, two-role selective swaps, outfit-preserving role edits, human → cat, and dog → cat, then kept the outputs that actually taught us something.

Short version: human → human works best, animal → animal can work, selective replace-one edits are the most practical multi-subject pattern, and human → animal is much less stable. Hands are usually the first thing to break.

Local runs only simple3 + reddit templates 4 to 6 second chunks Video examples included
Four MiniMax Ref2VA test frames showing human-to-human, replace-one-role, human-to-cat, and dog-to-cat examples
Four useful test buckets from the same local workflow: one-person human swap, replace-one-role dialog, human-to-cat, and dog-to-cat selective replacement.
Best overall

Single-person human → human replacement, especially when the source is centered, short, and clear.

Best selective edit

Two-role or two-subject clips where only one subject is replaced and the other stays untouched.

Most fragile area

Human → animal swaps, especially when hands, fingers, or small screen-within-screen details must also change.

Workflow rule

Chunk long clips into roughly 4 to 6 seconds, run each piece with the same prompt logic, then stitch.

The short scorecard

What actually worked in these Ref2VA tests

The easiest mistake is treating every swap type like the same task. It is not. The model behaves very differently depending on whether the job is identity replacement, role isolation, species change, or background replacement.

Fast verdicts

Six practical reads from the runs

Worked well
One person → one person

Best category overall. The cleanest results came from short, single-subject clips with readable motion and a clear replacement identity.

Original
Generated
Useful, but imperfect
Two people / two roles → replace one human

More stable than replacing both roles, but this result still had a real timing problem: parts of the dialog did not always feel synced to the right person.

Original
Generated
Worked with chunking
Animal → animal

Selective replacement can work, especially when the unaffected subject and props are explicitly frozen and the clip is split into 4 to 6 second chunks.

Original
Generated
Partial only
One person → one animal

Scene preservation can be decent, but arms and hands often survive, especially around object interaction.

Original
Generated
Weak
Two-role human dialog → cat replacement

Much weaker than human → human. The model can isolate a role, but it tends to drift into portrait-like cat shots or leave human remnants.

Original
Generated
Worked, with tradeoff
Two-role wardrobe preservation

Role-specific clothing can be preserved, and this keep-outfits case also held speaking timing better than Case 2, but stronger wardrobe preservation can start to weaken same-person identity across both roles.

Original
Generated
Prompt template choice

When to start from simple3 and when to switch to reddit

Template 1

simple3 is the safer start for one clear subject

Use the simple3 family when the job is fundamentally one subject in, one subject out.

  • Single-person dance or performance swap
  • One identity replacing another identity
  • Full-body cases where the motion path is already clean
simple3.prompt.txt is the base template, and simple3-girl-dance-hires.prompt.txt was one of the cleanest one-person human → human runs because it focused on the replacement subject's appearance and kept the scene simple.
Template 2

reddit is better when only one role or one subject should change

Use the reddit-style structure when the real challenge is not identity alone, but selective replacement.

  • Two-role dialog where only one speaker changes
  • Two animals where only one animal changes
  • Cases with important props that must remain intact
reddit.prompt.txt is the base template, and reddit-girl-first-role-only-4s.prompt.txt plus reddit-replace-right-dog-and-phone-dog-with-cat-4s.prompt.txt both worked because they explicitly named the subject to replace and the subject to preserve.
What we changed inside the templates

The prompt edits that mattered most

1. Say exactly which role changes and which role stays

  • Use language like Replace ONLY the first/light-shirt role or Do NOT replace the left-side husky puppy.
  • When this is missing, the model is more likely to either replace everything or mix identities.

2. Preserve pose and gesture by naming them

  • For dialog cases, writing the actual pose beats helped more than a vague “preserve pose” instruction.
  • Examples: open-palmed flourish, head tilt, clasped hands, raised-paw reacting gesture.

3. Scene integration must be explicit for animals

  • Without this, human → cat often collapses into a nice-looking portrait animal instead of a scene-consistent replacement.
  • Prompting for original room, framing, and “not a standalone studio portrait” helped, but did not solve hands fully.

4. Multiple references help identity, but not by magic

  • For human → human dance, a second reference improved consistency somewhat.
  • It was useful, but not a dramatic jump to perfect identity lock.

5. Tiny secondary screens are a weak point

  • Main-subject replacement can work while the tiny phone-screen subject stays ambiguous or half-preserved.
  • The model prioritizes the main body in the scene over the little screen-in-screen detail.
The biggest workflow lesson

One person → one person and two-subject selective replacement are the best starting points

If you want one simple starting rule, use this: first test the easiest solvable structure, not the most exciting idea. For Ref2VA, that means one clear subject or one clearly named replace-only target.

When the task becomes “replace both roles,” “replace the background only,” or “make a human become an animal while preserving hands and object interaction,” the failure rate goes up quickly.

Chunking and stitching

Why 4 to 6 second segments were the practical rule

  • 4-second runs were fast enough for prompt iteration.
  • 6-second runs were still manageable for long clips without overcommitting to a bad prompt.
  • Longer clips were more reliable when split into 4 to 6 second chunks and stitched after generation.
  • We kept the same prompt logic across chunks so the replacement rule stayed stable.
Bottom line

What I would recommend to creators after these tests

Start here

Single-person human swaps

If you are just checking whether Ref2VA is useful for your workflow, start with short one-person human → human tests.

Then move to

Replace-one-role or replace-one-subject jobs

These are a much better second step than “replace everything” ideas because they teach you whether role isolation is working.

Be careful with

Human → animal edits with visible hands or props

They can produce interesting outputs, but they are much more likely to leave hands, preserve the wrong limb structure, or drift into portrait-like compositions.