What we were actually testing
The source was a vertical one-person TikTok-style dance clip normalized to 24 fps. We split it into two legal MiniMax H3 chunks:
- Chunk 1: 158 frames, 6.583s
- Chunk 2: 175 frames, 7.292s
The hard requirement was not only “make this dancer become the blonde woman.” The hard requirement was:
For every frame where the woman appears: keep the same clothing keep the same body motion keep the same pose keep the same head orientation keep the same expression timing keep the same scene and camera The visible head must be <Subject 1> from the first frame to the last frame. Do not replace the clothing. Use <Picture 1> for all visible head identity. If there is any conflict: keep the body from <Video 1> keep the clothing from <Video 1> keep the pose from <Video 1>
What happened across the four tries
The custom reddit-style chunk1 prompt proved that identity transfer and scene carryover were possible, but it did a poor job preserving source outfit changes.
The stricter simple3-based chunk1 rerun moved closer to the intended “only replace the visible head” behavior, but later frames still drifted away from the original room and clothing.
Chunk2 stayed unstable. One prompt produced a few good late frames, while the base simple3 rerun was slightly better overall but still not good enough to call solved.
Original chunk on the left, generated result on the right
Case 1 · Chunk 1 · custom reddit prompt
Good identity transfer, decent framing, weak source-outfit fidelity.
This attempt showed that the model could place the blonde identity into the source dance scene. But it read more like “put the blonde woman into this choreography” than “keep every source cut's exact clothing and pose logic.”
Case 2 · Chunk 1 · simple3 preserve-clothing rerun
Better direction for head-only replacement, but still not stable enough.
This was the first try that felt closer to the exact prompt goal: keep the source body and scene, use the picture for visible head identity. The problem was that later rows still drifted into a plainer studio look, which took clothing fidelity down with it.
Case 3 · Chunk 2 · chunk-specific simple3 prompt
A few promising late frames, but early cuts collapsed hard.
The local read on this run was very uneven: rows 1–2 drifted into a studio-like background, rows 3–4 were mixed, and row 5 was the best-looking frame. That is not enough for a reusable workflow when the user cares about consistent clothing across cuts.
Case 4 · Chunk 2 · base simple3 rerun
Slightly better than Case 3, but still not a success.
This rerun held the blonde identity and the curtain/carpet environment a bit longer than Case 3, and the first row looked strong. But from row 2 onward, clothing started copying the wrong shot and later frames still drifted hard.
The rules this case taught us
simple3 for this structure
This was still a one-person → one-person dance replacement task. The clip had outfit complexity, but it was not a selective multi-subject edit. That kept simple3 as the right starting family.
Because the source changed clothes across cuts, the prompt needed an explicit conflict rule: keep body, clothing, pose, and scene from the video; use the picture only for visible head identity.
The model could often keep the blonde identity or keep part of the source scene, but holding the exact outfit for each cut across the whole chunk remained fragile.
Once the curtain-and-carpet room simplified into a cleaner studio-like backdrop, outfit fidelity usually fell apart in the same stretch.
Chunk2 contained enough additional wardrobe pressure that a prompt which looked directionally right on chunk1 still struggled later.
We kept audio in the workflow and verified stream presence, but the bigger blocker in this case was the visual inconsistency. The publishable problem here was not missing audio. It was unstable clothing and scene retention.
Our honest verdict
MiniMax Ref2VA could usually put the blonde identity into the same general dance clip, and sometimes preserve enough of the curtain-and-carpet scene to look promising for a few frames.
It did not reliably preserve the source clothing across all jump cuts. That was the deciding failure, because preserving the source wardrobe was the real goal of the test.