VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Tencent

Research official + media 2 src. ~1 min

A unified framework plus VWE-BENCH (2,616 assets, 323 seed worlds, 6,828 queries) and VibeWorlding-Gym (sandbox + rubric verifier) for training multimodal agents that infer intent, invoke 3D tools, and reflect on feedback. Even GPT-5.5 and Qwen3.8-Max score below 60%; RL-trained VibeWorlder-30B-A3B takes the best Pass@1 overall, beating closed-source models.

Why it matters

HF: 34 upvotes. First open benchmark plus training recipe showing open MLLMs can surpass frontier closed models on end-to-end 3D world construction.

Importance: 2/5

HF Daily 34 upvotes

Sources

official arXiv:2608.15265