#vision-language-action
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM research
- τ₀-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation Shanghai Innovation Institute research
- Xiaomi-Robotics-1: scaling vision-language-action models with 100K+ hours of real-world trajectories Xiaomi Robotics research