ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
Researchers have introduced ActCam, a novel zero-shot method for video generation that enables fine-grained, joint control over both character motion and camera trajectories. Designed for artistic applications, ActCam transfers actor movements from a source video into new scenes while allowing per-frame adjustment of intrinsic and extrinsic camera parameters. The system operates on pretrained image-to-video diffusion models by generating geometrically consistent pose and depth conditions. It employs a unique two-phase conditioning schedule: initial denoising steps utilize both pose and sparse depth to establish scene structure, followed by pose-only guidance to refine high-frequency details without over-constraining the output. Evaluations across diverse benchmarks demonstrate that ActCam significantly improves camera adherence and motion fidelity compared to existing pose-only or combined methods, particularly during large viewpoint changes. Human evaluations further confirm its superiority in maintaining visual consistency. This approach highlights the effectiveness of careful camera-consistent conditioning and staged guidance in achieving robust video generation without additional model training, offering a powerful tool for creators seeking precise cinematic control.
Wire timeline
ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video Generation
Researchers have introduced ActCam, a novel zero-shot method for video generation that enables fine-grained, joint control over both character motion and camera trajectories. Designed for artistic applications, ActCam transfers actor movements from a source video into new scenes while allowing per-frame adjustment of intrinsic and extrinsic camera parameters. The system operates on pretrained image-to-video diffusion models by generating geometrically consistent pose and depth conditions. It employs a unique two-phase conditioning schedule: initial denoising steps utilize both pose and sparse depth to establish scene structure, followed by pose-only guidance to refine high-frequency details without over-constraining the output. Evaluations across diverse benchmarks demonstrate that ActCam significantly improves camera adherence and motion fidelity compared to existing pose-only or combined methods, particularly during large viewpoint changes. Human evaluations further confirm its superiority in maintaining visual consistency. This approach highlights the effectiveness of careful camera-consistent conditioning and staged guidance in achieving robust video generation without additional model training, offering a powerful tool for creators seeking precise cinematic control.
cs.AI updates on arXiv.org