SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen, Ying Shen, Kiet A. Nguyen, Adheesh Sunil Juvekar, Ismini Lourentzou
- Published
- Oct 1, 2026 — 17:59 UTC
Problem
The paper addresses the fragmentation of continuous surfaces in high-resolution 3D generation, which leads to increased generation costs and weakened topological consistency. This issue is particularly relevant in the context of generating complex 3D structures where maintaining surface integrity is crucial. The work is presented as a preprint and has not undergone peer review.
Method
The proposed framework, SILSA (Sliding-Window Slice Latents), utilizes a compact representation of sliding-window slice latents. The method involves a fixed set of overlapping slices along three canonical axes, which allows for effective encoding and decoding of 3D structures. The encoding process employs a Slice Variational Autoencoder (VAE) that encodes oriented surface samples into multi-axis slice latents. For decoding, a sparse volumetric decoder reconstructs these slice latents into a coherent 3D representation. Additionally, a Volumetric Anchor Lattice coordinates the directional slice streams, enhancing the model's ability to maintain topological consistency. The framework incorporates slice-level topology supervision, which matches persistence diagrams and aligns Betti transitions across neighboring slices, ensuring that the generated 3D models preserve their topological features.
Results
The results demonstrate significant improvements over the strongest baseline models:
- PSNR Improvement: 8.7%
- Coverage Improvement: 5.96 absolute points
- Betti Error Reduction: 9.2%
- Token Reduction: 70.0% fewer tokens compared to the next-most compact baseline
- Training Memory Reduction: 40.4% compared to previous methods
- Inference Time Reduction: 58.5% compared to previous methods These metrics indicate that SILSA not only enhances the quality of the generated 3D models but also optimizes resource usage during training and inference.
Limitations
The authors do not report any limitations in their work, suggesting that the proposed method effectively addresses the identified problems without notable drawbacks.
Why it matters
The implications of this work are significant for downstream applications in 3D modeling, computer graphics, and virtual reality, where high-resolution and topologically consistent models are essential. By reducing computational costs and improving the quality of generated 3D structures, SILSA could facilitate more efficient workflows in industries reliant on 3D generation, such as gaming, simulation, and design.
By Turing Wire Research Desk · Oct 1, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
