Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene… — HuggingFace Daily — TG.ME

"Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling" by Minghan Qin , Yuang Wang , Xiuyu Yang , Yushi Long , Yujian Zhang , Ruihuan Wang , Kai Ye , Yangang Zhang , Hang Li

TLDR:
Lucida is a novel system that enhances indoor scene reconstruction by distributing pipeline requirements across parsing, asset generation, and VLM-guided placement to create high-quality editable replicas from cluttered captures. Compared to existing methods that struggle with accurate geometry and occlusions in cluttered captures, Lucida proposes a more flexible approach that considers the limitations of real captures. It parses videos into scenes with multi-view evidence, generates assets from this evidence, and places them using a VLM policy called GizmoAct, which leads to precise alignment. Through its innovative approach, Lucida significantly improves performance metrics, showcasing a 69% increase in mAP over Boxer, a boost in [email protected] from 57.8% to 83.4%, and an enhancement of scene F-Score from 0.794 to 0.924.

Read Paper / Blog
huggingface.co
Paper page - Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Join the discussion on this paper page
September 2, 2026 70 1