#attention
- A Systematic Analysis of Hybrid Linear Attention: 72-Model Study ByteDance Seed research
- The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models research
- MiniMax Sparse Attention: 28× Compute Reduction at 1M-Token Context with No Quality Loss MiniMax research
- FlashMorph: Data-Driven Hybrid Attention Layer Placement via Learnable Gates ByteDance Seed research