Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, Wu Liu, Xi Peng, Chun Jian Ho, Hongyuan Zhu
- Published
- Sep 28, 2026 — 17:55 UTC
Problem
The paper identifies a gap in the effective initialization of linear Vision Transformers (ViTs) using pre-trained weights from Softmax ViTs. This issue is particularly relevant as the performance of linear ViTs can be significantly influenced by how their weights are initialized, yet existing methods do not leverage the potential of pre-trained models effectively. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose a novel approach for initializing linear ViTs, focusing on the transfer of weights from Softmax ViTs. The architecture utilized is based on Linear Vision Transformers. The core technical contribution includes a proper loss design aimed at distilling the attention routing behavior from Softmax ViTs to linear ViTs. The transfer mechanism is twofold: MLP weights can be directly copied in an operator-agnostic manner, while attention weights are operator-specific and do not transfer effectively. The datasets used for training are varied, but specific sizes and details are not disclosed. Training compute requirements are also not specified, which may limit reproducibility.
Results
The results indicate that linear ViTs can match and even surpass the performance of Softmax ViTs when employing copied MLPs and distilled attention. However, the available text does not report quantitative results or specific performance metrics against named benchmarks, which would provide a clearer context for the improvements claimed.
Limitations
The authors note that attention weights do not transfer effectively between Softmax and linear ViTs, which poses a challenge for the proposed initialization method. Additionally, the lack of specific model sizes and dataset details limits the applicability of the findings and may hinder further research based on this work. The absence of quantitative results also restricts the ability to evaluate the performance improvements in a rigorous manner.
Why it matters
This research has significant implications for the development of linear ViTs, particularly in enhancing their initialization strategies. By demonstrating that linear ViTs can leverage pre-trained weights effectively, the findings could lead to more efficient training processes and improved performance in various vision tasks. This work opens avenues for further exploration into weight transfer mechanisms and their impact on model performance, potentially influencing future architectures and training methodologies in the field of computer vision.
By Callan Zhang · Sep 28, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
