Beyond Data Scaling: Representation-Centric Continued Pre-training… — HuggingFace Daily — TG.ME

"Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models" by Senqiao Yang , Chengyao Wang , Yuxin Chen , Zixuan Wang , Longxiang Tang , Haokun Gui , Jinhui Ye , Changsheng Lu , Xiaoyang Wu , Mingkang Zhu , Pengguang Chen , Shu Liu , Zhuotao Tian , Hengshuang Zhao , Bei Yu , Jiaya Jia

TLDR:
The VLAct model is designed to enhance the performance of Vision-Language-Action models by pre-training them on a diverse range of robot data to retain vision-language priors and shared action semantics. The model achieves strong results in simulations and with limited computational resources across various embodiments. Scaling robot data plays a crucial role in developing generalist VLA models; however, collecting robot trajectories is more challenging than gathering image-text data due to cost and coverage limitations. VLAct focuses on improving representation quality by transforming limited trajectories into transferable visual-action knowledge. The model is trained on heterogeneous robot data, encouraging shared action semantics across embodiments. Results show that VLAct outperforms existing VLA systems in different scenarios and embodiments, even surpassing them with limited downstream trajectories and computational resources. The research demonstrates that continuing pre-training with a representation-centric approach can significantly enhance VLA model performance within constrained computational budgets.

Read Paper / Blog
huggingface.co
Paper page - Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
Join the discussion on this paper page
👍1
September 1, 2026 62