PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu
- Published
- Sep 15, 2026 — 17:55 UTC
Problem
Existing controllable video generation methods either require predefined control schedules or rely on pixel-space signals for object positioning, which do not effectively leverage physical dynamics. This paper addresses these limitations by introducing a novel approach that integrates structured scene memory and fine-grained motion control, enabling more realistic and controllable video synthesis. The work is presented as a preprint and has not undergone peer review.
Method
The proposed model, PhysStream, is an autoregressive architecture designed for physics-grounded image-to-video synthesis. Key components include:
- Structured Scene Memory: This component utilizes positional maps and object tracking maps that are derived online from previously generated frames, allowing the model to maintain context and continuity in the generated video.
- Motion Control: The model employs sparse velocity-increment signals that encode physical quantities, facilitating precise control over object motion.
The training process consists of two stages:
- Stage 1: A bidirectional model is finetuned with motion-control conditioning to establish a baseline for motion dynamics.
- Stage 2: A causal autoregressive model is trained, incorporating the structured scene memory to enhance the generation process.
Details regarding the specific data used for training and the computational resources required are not disclosed in the paper.
Results
The performance of PhysStream is evaluated against strong baselines using several metrics:
- Motion Distribution Distance (FVMD): Achieved a reduction of 33% compared to the strongest baselines, indicating improved motion realism.
- Trajectory Error: Reduced by 12% against the strongest baselines, demonstrating enhanced accuracy in object movement.
- Human Evaluator Preference: In comparisons conducted in real-world scenarios, PhysStream was preferred in over 85% of cases against prior methods, highlighting its effectiveness in generating visually appealing and coherent videos.
Limitations
The authors do not report any limitations in the study. However, the lack of specified training data and compute resources may hinder reproducibility and scalability assessments.
Why it matters
The implications of this work are significant for downstream applications in video generation, particularly in fields requiring high levels of control and realism, such as virtual reality, gaming, and simulation. By addressing the limitations of existing methods, PhysStream paves the way for more sophisticated and physically grounded video synthesis techniques, potentially influencing future research directions in controllable generative models.
By Callan Zhang · Sep 15, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
