#vision-language
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Images Technion research
- S²VOPD: self-supervised visual on-policy distillation lifts Qwen3.5-4B from 70.7% to 77.4% on six fine-grained perception benchmarks UC San Diego research
- HumanCLAW: Can Vision-Language Models Act Through a Body? Meta Research research
- DeepSeek releases experimental multimodal V4-Flash-Vision-Exp DeepSeek models-llm
- JoyAI-VL-Interaction: Open-Source 8B Real-Time VLM with Autonomous Turn-Taking JD.com research
- MemLens: Benchmark for Multimodal Long-Term Memory in Vision-Language Models NVIDIA research
- GitHub Copilot Vision Generally Available for All Plan Tiers GitHub tools
- PerceptionRubrics: Atomic Rubric Evaluation Reveals 8% Perception Gap Between Open and Closed Models research
- ABot-N1: Visual Language Navigation Foundation Model with Slow-Fast Architecture Alibaba research
- Qwen ships Qwen3.8-27B and FP8 variant as Apache-2.0 open-weights vision-language model Qwen models-llm
- Astra: RL-Trained VLM Queries World Simulator for Spatial Reasoning research
- Tencent releases HY-Embodied-0.5-X update for embodied agents Tencent models-llm
- Visual Contrastive Self-Distillation University of Maryland research
- GST-Bench tests whether VLMs develop global spatial awareness from video ByteDance Seed research