| |
Gemma 4 12B: A unified, encoder-free multimodal model
Google DeepMind has introduced Gemma 4 12B, a unified multimodal model featuring an encoder-free architecture that processes vision and audio inputs directly into its language model backbone, enabling advanced reasoning and agentic capabilities on standard laptops with just 16GB of RAM. The model delivers performance comparable to larger 26B models while requiring less than half the memory footprint and is the first mid-sized Gemma model to natively support audio inputs. Released under an Apache 2.0 license, Gemma 4 12B is designed to bring high-performance multimodal AI directly to consumer hardware without sacrificing speed or reasoning ability.
Read Full Article →
← More Tech news