TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters The paper introduces Tokenformer. The architecture leverages the attention mechanism to facilitate not only inter-token computations but also interactions between tokens and model parameters. The authors replace all linear projection layers in the Transformer with Pattention layers, allowing for efficient incremental scaling without the need for retraining from scratch. Future work: - Extending the Mixture-of-Experts Paradigm - Advancing Parameter-Efficient Tuning - Integrating Vision and Language Models - Device-Cloud Collaboration - Enhancing Model Interpretability Code: https://github.com/Haiyang-W/TokenFormer
November 8, 2024 596 8