GenFirst: Generation Before Reconstruction for Stable End-to-End… — HuggingFace Daily — TG.ME

"GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling" by Guangting Zheng , Yiyuan Zhang , Tao Yang , Yunpeng Chen , Rui Zhu , Jiajun Deng , Yanyong Zhang

TLDR:
The text discusses the challenges and solutions in training latent generative models directly from end to end to improve image synthesis and multimodal generation. Typically, these models involve a two-stage process of training a variational autoencoder for reconstruction and then a generative model separately. The text proposes a novel approach that involves training both models simultaneously to address the limitations of reconstruction-optimized latents inhibiting generation. By preserving entropy and understanding the asymmetric dynamics of reconstruction and generation, the authors introduce a new strategy, GenFirst, which enables successful end-to-end training without latent collapse. This strategy involves shaping the latent space first for generation before reinforcing it for detailed reconstruction. The approach is validated across various generative priors and modalities, showcasing improved performance in tasks like text-to-image generation and shared visual-latent representation learning.

Read Paper / Blog
huggingface.co
Paper page - GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
Join the discussion on this paper page
September 2, 2026 45