Project Page

UniMoCa

Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

Liming Tan*1, Ye Chen*1, Hao Zhang3, Lirong Qian1, Feifei Li2, Bingbing Ni1,2

1 Shanghai Jiao Tong University
2 USC-SJTU Institute of Cultural and Creative Industry
3 Monash University

* Equal contribution. Corresponding author.

Demo Code Coming Soon Data Coming Soon

Abstract

Faithful Human Video Generation with Unified Visual Control

Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences for motion control while using camera embeddings for camera control, forcing video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings.

UniMoCa is a representation-driven framework that unifies motion and camera controls in visual space. At its core is Motion-Camera Visual Proxy (MCVP), an identity-neutral representation that converts 3D human motion and camera trajectories extracted from driving videos into a shared visual proxy with explicit camera trajectory markers. Experiments based on Wan2.2 I2V show strong gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity.

MCVP Unified visual proxy
80K video-proxy pairs
120h training video data
Wan2.2 I2V backbone

Motivation

Motion and Camera Controls Belong in the Same Visual Space

Prior methods mix visual motion cues with geometric camera codes, which makes it hard for the model to decide whether an observed change comes from the body or the camera. UniMoCa turns both factors into synchronized, distinguishable visual cues.

Motivation diagram comparing heterogeneous controls with UniMoCa MCVP unified visual proxy.
UniMoCa replaces heterogeneous visual-parametric controls with MCVP, reducing motion-camera attribution ambiguity.

Method

Motion-Camera Visual Proxy Guided Generation

UniMoCa recovers camera-view human motion and camera trajectories from driving videos, transforms them into a global-view motion sequence, renders human and camera visual factors into MCVP, and feeds proxy tokens into the video diffusion transformer with Shifted RoPE.

Overview of UniMoCa showing MCVP construction and MCVP-guided video generation.
Overview of MCVP construction and proxy-token conditioning for faithful motion and camera alignment.

Motion Control

Motion Control Results

Each demo uses the composited motion video prepared for the project page, showing the control signal and generated result in a single clip.

Boxing

Running

Skateboarding

Hurdling

Two-person Dance

Jogging

Football

Action Scene

Ballet

Skating

Skiing

Camera Control

Consistent Motion Under Camera Editing

Camera-control demos group the same action across available view changes. UniMoCa preserves motion and subject layout while following static, lateral, upward, and dynamic camera directions.

Fight

Static
Left
Right
Up

Slap

Static
Left
Right
Up

Interview

Static
Left
Right
Up

Dance

Static
Left
Right
Dynamic

Kick

Static
Left
Right
Up

Kungfu

Static
Left
Right
Up