For a while, we thought Mountain View had “lost its way” with its AI development, but it seems the turmoil is behind us. Through a post on its official blog, Google announced Gemini, a new multimodal model that aims to compete with GPT-4, the current system in the paid version of ChatGPT. Google describes Gemini as the most powerful model it has created so far, while also acknowledging the need for more flexibility by confirming the existence of three builds or “sizes”, one of which is already available via Bard.
Every AI model has its strengths and weaknesses, but one thing is certain: they must do much more than they do today. If we throw a stone in any direction, it's almost impossible not to hit a chatbot with it; however, the public needs to work with other content beyond text. So we arrive at the stage of “multimodal” models. Video, audio, code, and images become accessible to these platforms, and in Google's specific case, its answer is Gemini.
Gemini, Google's Multimodal Model That Will Compete with GPT-4
Google explains that previous examples of multimodality were simply different exclusive models (text-only, image-only, audio-only) joined inefficiently in secondary processing stages. Gemini doesn't use that strategy; instead, it was designed from scratch to incorporate multimodality into its structure.
Another essential aspect of Gemini is that it will be available in three modes or sizes: Gemini Ultra for high-complexity tasks that demand a lot of hardware, Gemini Pro aimed at general use, and Gemini Nano, designed for mobile devices and on-device execution.
For now, only Gemini Pro is accessible, courtesy of a build specially optimized for Google Bard. However, users in Europe and the UK will have to wait a bit longer (regulations, regulations). Another restriction is language: to explore Gemini/Bard's virtues, you have to speak English.
As expected, Google shared several benchmarks that place Gemini Ultra above or at the same level as GPT-4, but this doesn't change anything for the average user. In fact, there's a contradiction in presenting Gemini as a general solution while highlighting its performance through extraordinarily specific benchmarks.
At the same time, the underlying problems remain intact. AI still invents answers or “deliriums” in their sessions, and Gemini is not immune to this, but Google's bet on its multimodality is very big, and we'll surely see more about it in the near future.
Official announcement: Click here