6 articles about multimodal
Google introduces agentic video understanding in Gemini, letting models navigate, reason about, and answer questions on long videos — from content review to sports analysis.
DeepSeek quietly ships deepseek-v4-flash-vision-exp, an experimental vision variant of its V4-Flash model that accepts image input, alongside V4-Pro and V4-Flash point updates.
New research introduces Puffin-World, a unified multimodal model that works with native 3D world states — advancing AI's ability to understand and generate spatial environments.
Google's multimodal video generation model exits beta with 1080p output, 15-second clips, and 40% faster generation — now production-ready for developers.
MiniMax launches H3, a general-purpose multimodal generation model that unifies text, image, video, and audio in one context — 15-second 2K video with native stereo sound, with open weights promised within days.
Google launched Gemini Omni Flash, the first model in the Omni family, at I/O 2026 — a multimodal video generation and editing system that accepts text, images, audio, and video as input.