"LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes" by Chuyan Chen , Haoxing Chen , Kun Chen , Zhenglin Cheng , Long Cui , Ruishan Fang , Zhangxuan Gu , Zhicheng Huang , Zhenzhong Lan , Yuanting Lei , Haoquan Li , Jianguo Li , Rongchuan Li , Sidu Li , Tao Lin , Deyuan Liu , Jiacheng Liu , Lin Liu , Yuxuan Lou , Zhisheng Lu , Yuxin Ma , Shuheng Shen
TLDR:
LLaDA-Image is a cutting-edge framework that combines a 6B Diffusion Transformer and a vision-language module to produce photorealistic images with precise editing capabilities. The model is trained using image-only pre-training and a specialized optimizer to achieve state-of-the-art results in open-source benchmarks. By leveraging a strong visual generative prior and innovative optimization techniques, LLaDA-Image generates high-quality images while following detailed editing instructions accurately. Additionally, a faster variant, LLaDA-Image-Turbo, allows for efficient inference in just 2-4 sampling steps. The model sets new benchmarks in both English and Chinese tracks and is made available to support future research in generative models.
Read Paper / Blog
1September 4, 2026 52