#multimodal

6 articles about multimodal

Gemini Gains Agentic Video Understanding
AI Products Sep 12, 2026

Gemini Gains Agentic Video Understanding

Google introduces agentic video understanding in Gemini, letting models navigate, reason about, and answer questions on long videos — from content review to sports analysis.

DeepSeek V4-Flash Gets Vision — Experimental Image Input Comes to the Budget Model
AI Models Sep 3, 2026

DeepSeek V4-Flash Gets Vision — Experimental Image Input Comes to the Budget Model

DeepSeek quietly ships deepseek-v4-flash-vision-exp, an experimental vision variant of its V4-Flash model that accepts image input, alongside V4-Pro and V4-Flash point updates.

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
AI Research Sep 2, 2026

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

New research introduces Puffin-World, a unified multimodal model that works with native 3D world states — advancing AI's ability to understand and generate spatial environments.

Google Gemini Omni Flash Leaves Beta — Video Generation Goes GA
AI Products Aug 27, 2026

Google Gemini Omni Flash Leaves Beta — Video Generation Goes GA

Google's multimodal video generation model exits beta with 1080p output, 15-second clips, and 40% faster generation — now production-ready for developers.

MiniMax H3: One Open Model for Every Modality — 2K Video With Stereo Sound at a Third the Price
Open Source AI Jul 31, 2026

MiniMax H3: One Open Model for Every Modality — 2K Video With Stereo Sound at a Third the Price

MiniMax launches H3, a general-purpose multimodal generation model that unifies text, image, video, and audio in one context — 15-second 2K video with native stereo sound, with open weights promised within days.

Google Gemini Omni Flash Launches: The 'Nano Banana for Video' Creates and Edits Video From Any Input
AI Tools May 22, 2026

Google Gemini Omni Flash Launches: The 'Nano Banana for Video' Creates and Edits Video From Any Input

Google launched Gemini Omni Flash, the first model in the Omni family, at I/O 2026 — a multimodal video generation and editing system that accepts text, images, audio, and video as input.