#video-understanding
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding MCG-NJU research
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Microsoft research
- GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? ByteDance Seed research