ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
Four-stage progressive training (domain adaptation -> teacher-forced causal -> causal consistency distillation -> on-policy distribution matching) converts a bidirectional video generator into a 1/2/4-step causal world model. Demonstrated on Minecraft and gamepad-controlled FPS with dual-path deployment for interactive vs replay use.
Why it matters
HuggingFace Daily Papers: 95 upvotes. First framework (per authors) to support low-latency causal game control from a diffusion backbone at sub-4-step inference.
Importance: 2/5
HF Daily Paper with 95 upvotes (just below the 100-upvote importance bump)
Sources
official
arXiv abstract page
official
HF Daily Papers entry