Project Page
UniMoCa
Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
Abstract
Faithful Human Video Generation with Unified Visual Control
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences for motion control while using camera embeddings for camera control, forcing video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings.
UniMoCa is a representation-driven framework that unifies motion and camera controls in visual space. At its core is Motion-Camera Visual Proxy (MCVP), an identity-neutral representation that converts 3D human motion and camera trajectories extracted from driving videos into a shared visual proxy with explicit camera trajectory markers. Experiments based on Wan2.2 I2V show strong gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity.
Motivation
Motion and Camera Controls Belong in the Same Visual Space
Prior methods mix visual motion cues with geometric camera codes, which makes it hard for the model to decide whether an observed change comes from the body or the camera. UniMoCa turns both factors into synchronized, distinguishable visual cues.
Method
Motion-Camera Visual Proxy Guided Generation
UniMoCa recovers camera-view human motion and camera trajectories from driving videos, transforms them into a global-view motion sequence, renders human and camera visual factors into MCVP, and feeds proxy tokens into the video diffusion transformer with Shifted RoPE.
Motion Control
Motion Control Results
Each demo uses the composited motion video prepared for the project page, showing the control signal and generated result in a single clip.
Running
Skateboarding
Hurdling
Two-person Dance
Jogging
Football
Action Scene
Ballet
Skating
Skiing
Camera Control
Consistent Motion Under Camera Editing
Camera-control demos group the same action across available view changes. UniMoCa preserves motion and subject layout while following static, lateral, upward, and dynamic camera directions.