2 articles about speech-to-text
Google releases Gemini 3.5 Transcribe to general availability with real-time speech-to-text across 85 languages, word-level timestamps, and speaker diarization.
OpenAI introduces three new audio models in the API: GPT-Realtime-2 for conversational reasoning, GPT-Realtime-Translate for live translation, and GPT-Realtime-Whisper for streaming transcription.