WorldSculpt: Generating Compositional Worlds from Grounded Videos" by… — HuggingFace Daily — TG.ME

"WorldSculpt: Generating Compositional Worlds from Grounded Videos" by Muyao Niu , Jixuan He , Ruihan Yu , Lian Fu , Yonghao Yu , Zheng-Hui Huang , Yifan Zhan , Fengbo Lan , Yongtao Ge , Yinqiang Zheng , Kaipeng Zhang , Zhixiang Wang

TLDR:
The text discusses the complex task of generating a detailed 3D representation of cluttered scenes with many objects using a single-object 3D generative prior adapted to multi-view observations. The aim is to reconstruct scenes as collections of individual object meshes within a shared world framework, crucial for applications like gaming, AR/VR, simulation, and robotics. Traditional methods struggle with scenes where heavy occlusion and limited views hinder complete reconstruction. By applying a transformative approach called Pixal3D with a multi-view pathway, the model excels in generating complex scenes with hundreds of objects without specific scene-level training, showcasing scalability. The paper introduces UE-MeshyScene, a benchmark dataset, and demonstrates superior performance across various evaluations, especially in scenarios with increased complexity and occlusion. The method's versatility is showcased by converting other 3D environments into compositional mesh scenes successfully.

Read Paper / Blog
huggingface.co
Paper page - WorldSculpt: Generating Compositional Worlds from Grounded Videos
Join the discussion on this paper page
September 7, 2026 58