| |
Google has introduced Gemini 3.5 Transcribe, its most advanced speech-to-text model designed for precise, real-time transcription that handles background noise, complex jargon, and natural speech patterns like self-corrections and filler words. The model is now available to developers through the Gemini API, Google AI Studio, and Gemini Enterprise Agent Platform via two APIs for real-time streaming and pre-recorded audio processing. With a word error rate of 4.0% for streaming and 2.6% for non-streaming, the model supports over 85 languages and includes features like custom vocabulary recognition and function calling for delegating complex tasks.
Read Full Article →
← More Tech news