Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

Research official 2 src. ~1 min

Introduces a 1,253-example, expert-annotated benchmark across 60 subjects in four scientific disciplines for testing whether video generation models get scientific facts and causal dynamics right, not just visual realism. Finds models vary widely on scientific correctness with a pronounced gap between proprietary and open-source systems.

Why it matters

21 upvotes on HuggingFace Daily Papers; addresses a gap in video-gen evaluation that current benchmarks (visual quality only) miss.

Importance: 2/5

Notable paper (21 HF Daily Papers upvotes, below the 100-upvote bump threshold).

Sources