You give it two things: a driving video and one frame from that same video with the character repainted, in any image editor or with an image model. The model carries the edit across every frame, so a person doing a kick becomes a corgi, a robot, or a clay figure doing the same kick, on the same background.
What it does not need is the point. There is no pose skeleton, no mask, no face crop, and no text prompt. Because the reference frame comes from the clip itself, its pose and framing already match the footage. That is also why it holds up under fast motion: whipping head turns and full jumps track frame for frame instead of smearing. And since nothing in the loop assumes a human body, it animates animals, an airliner with wings bound to the actor's arms, and several characters in one shot.
It is a finetune of MiniMax-H3's video model, distilled so a five-second shot renders in about 26 seconds on a data-center GPU, roughly six times faster than Wan2.2-Animate.