I must admit that I have abandoned AI image generation. The latest models demand all the VRAM in the world, deliver very limited results, or sit behind a paywall. However, Google reminds me that it's possible to access the free version of Gemini 2.0 Flash and use its integrated edition of the Imagen 3 model to generate images while maintaining a conversation. How well does it work?
Yes, the web has been invaded by AI-generated images. No, they're not very good, to say the least. In fact, the vast majority is, in a word, horrible. I admit I've used them in the past, but I stopped because the quality factor simply isn't there. At the same time, it's striking that major brands insist on betting on such mediocre visual results and ignore the criticism (I guess saving a designer's salary is more important).
But its evolution continues, or at least that's what the hype suggests. Google announced the opening of an experimental version of Gemini 2.0 Flash for its AI Studio platform, which includes multimodal support, advanced reasoning, and better natural language understanding. By far, the most notable feature of that package is "conversational editing" of an image (something we've already explored in the past), however, today I'm interested in seeing the current state of image generation under the conventional version of Gemini. I opened a new chat and shared some ideas...
Generating images in a chat, ft. Google Gemini 2.0 Flash
The only "advantage" I decided to give Gemini was starting and maintaining the conversation in English, to minimize any misinterpretation. First I described a bus stop, on a rainy night, lit by a streetlight, with city lights in the distance, covering the landscape. The initial results were very promising, but the generator started making mistakes and ignoring important aspects as I requested more details. In one image, all I asked was to rotate the streetlight 90 degrees to change the direction of the light... and it ended up replacing everything with a new image.
What about generating people? Gemini is without a doubt much more confrontational. Black spandex suits for our space heroines are not a problem, but when requesting a bikini or a swimsuit, clashes with its filter became more frequent. It also suffered the same problem as the bus stop: There came a point where it ignored basic elements of the description, like eye color.
To finish, some food: The first ice cream bowl was actually a dessert bowl, although very well done. Then it correctly changed it to three scoops of ice cream (two chocolate, one vanilla). But when I asked to improve the texture, it decided to insert a fourth scoop of ice cream... and never remove it.
So, I leave this session of Google Gemini 2.0 Flash having reached the daily limit of generated images, and feeling exactly what I imagined before starting: The disappointment of 'almost, but not'. I guess I can click the "Redo" command and surrender completely to the algorithmic gods, hoping that the next seed number is more favorable... but at the end of the day, nobody wants that. The "Eureka moment" still feels far away, and if you allow the pessimism, it's likely it will never come.
Access Gemini: Click here