Generative Cinematographer: Composing Camera and Object Motion in 3D
Jiahan Zhang, Chaohao Yang, Namitha Guruprasad, Vivekjyoti Banerjee, Trong-Tung Nguyen, Alan Yuille, Anand Bhattad
- Published
- Oct 1, 2026 — 17:58 UTC
Problem
The paper addresses the ambiguity in 2D motion trajectories when representing 3D motion, which is a significant gap in the existing literature. This issue complicates the generation of coherent 3D scenes from 2D inputs. The work is presented as a preprint and has not undergone peer review.
Method
The authors introduce the Generative Cinematographer (GenCine), which takes a single image as input and outputs an editable 3D scene scaffold. The system employs a control mechanism that includes:
- Camera Path Specification: Users can define the camera trajectory.
- Local 3D Motion Handles: These are provided for foreground regions to manipulate motion directly.
- Piecewise-Rigid Approximation: This technique is used to handle non-rigid motion effectively.
The projection of controls into guidance maps allows for the recording of controlled regions in each frame, with a fixed color assignment for handles across frames. The system encodes 3D positions in a common world coordinate system, facilitating coherent scene generation.
For training, GenCine utilizes controls derived from real video motion alongside ground-truth geometry and trajectories sourced from synthetic videos. The training architecture includes a lightweight guidance branch, LoRA adapters, and a pretrained Wan model, optimizing the system for performance and efficiency.
Results
The results indicate significant improvements in several areas:
- Camera-relative Motion Consistency: Enhanced compared to prior methods.
- Geometric Consistency Under Viewpoint Changes: Stronger performance than baseline approaches.
- Controllability Across Diverse Real-World Scenes: The system demonstrates enhanced controllability compared to previous systems, allowing for more flexible scene manipulation.
The available text does not report quantitative results.
Limitations
The authors do not report any limitations in their work. However, the absence of quantitative results may limit the ability to fully assess the performance improvements claimed.
Why it matters
The implications of this work are significant for downstream applications in computer graphics, virtual reality, and augmented reality, where the ability to generate and manipulate 3D scenes from 2D inputs can enhance user experience and creative possibilities. The advancements in motion consistency and controllability could lead to more intuitive tools for artists and developers in these fields.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
