#audio
- MiniMax Launches H3, an Omni-Modal Model Generating 2K Video With Native Stereo Audio MiniMax video
- Alibaba launches Wan 3.0 video model with 30-second single-pass generation and native audio Alibaba (Tongyi Lab) video
- ElevenLabs Music v2: Mid-Track Genre Switching, Inpainting, and Commercial Clearance ElevenLabs audio
- ElevenLabs Music v2 API Goes Live with Genre-Switching and Inpainting ElevenLabs audio
- ByteDance Launches Seed-Audio 1.0: Unified Speech, Music, and Ambient Sound Generation ByteDance audio
- EchoWM: Open and Enterable Omnimodal World Models JD.com research
- TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming Tencent research
- Google Brings Veo 3.1 Audio to All Flow Editing Tools, Adds Insert and Remove Google DeepMind video
- Wan-Streamer v0.1: End-to-End Real-Time Interactive Foundation Model Under 550ms Latency Wan-AI research
- Suno Launches Advanced Stem Separation with Per-Instrument Extraction Suno audio
- OpenAI adds SynthID watermarking to GPT-Live audio with API verification OpenAI tools
- Stable Audio 3.0 gains a DAW plugin and multitrack web workflow Stability AI audio