📈 Generative Modeling — Scaling, Multimodality & End-to-End Generation
✅ This Week's Presentation:
🔹 Title: Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
🔸 Presenter: Amir Qeysarbeigi
🌀 Abstract:
Modern generative models typically handle multimodal distributions by factorizing the generation process into multiple steps, as in autoregressive and diffusion models. While this enables high-quality generation, it creates a mismatch between training and inference and prevents fully end-to-end generation. In this work, we introduce Explorative Modeling (XM), a new paradigm that instead factorizes the training process by exploring multiple candidate generations and training on the best-matching one. This exploration increases generative expressivity, allowing models to capture more modes of multimodal distributions without relying solely on generation factorization.
The paper demonstrates that exploration acts as a third pretraining scaling axis, alongside model parameters and data, improving efficiency across image, video, and language generation. Increasing exploration improves FLOP, sample, and parameter efficiency, with gains that become larger as models and datasets scale. Furthermore, by moving the burden of multimodality from inference-time generation steps to training-time exploration, Explorative Modeling enables end-to-end generative models that can achieve performance comparable to diffusion-based approaches with dramatically fewer inference steps.
📄 Article: *Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation*
Session Details:
We will first review the limitations of conventional reconstructive generative models and introduce the concept of generative expressivity as a fundamental bottleneck in multimodal generation. Then, we will explore how Explorative Modeling replaces generation factorization with training-time exploration, including the Forward and Reverse XM formulations. Finally, we will examine how exploration serves as a new scaling axis, improves efficiency across multiple modalities, and enables end-to-end generation with substantially fewer inference steps.
- 📅 Date: Monday (دوشنبه)
- 🕒 Time: 17:00 - 18:00
- 🌐 **Location: https://vc.sharif.edu/rohban** (Online only)
We look forward to your participation! ✌️