Here at NeoTeo we have been following Microsoft's efforts in artificial intelligence. From age detection and emotion recognition to a chatbot that became psychotic and racist in less than 24 hours, the people at Redmond are definitely determined to improve their technology with multiple online projects. Now it's CaptionBot's turn, a system where you just upload a photograph and let the algorithm generate an automatic description.
Recently I talked about Seeing AI, an AI platform whose main function is to assist blind people. One of the most critical points of this technology is its ability to describe the user's surroundings accurately. In the presentation video we see a test when it identifies a young man doing a skateboard trick, but that's just a taste of its potential. Needless to say, all AI developments (whether they belong to Microsoft or not) have a huge amount of work ahead. The Redmond giant learned a thing or two with Tay, its ex-chatbot lover of fascism and incest, especially that it needs good teachers instead of "the Web", and decided to return to more controlled experiments, such as CaptionBot.
How CaptionBot works
Basically, all you can do in CaptionBot is upload a photo or paste a URL that leads to a compatible image, and let the AI generate an automatic description. CaptionBot announces from the start that it will keep the photo for a while to improve its skills, but that doesn't include personal information. During my tests I've noticed that some of its descriptions can be very superficial (although technically correct), but it's not hard to steer it into making several mistakes in a row. When the identification parameters are not ideal, CaptionBot will introduce phrases like "I’m not really confident" into its description.
In summary: Far... but getting closer. If we combine all the detection and recognition platforms that Microsoft is developing with what HoloLens represents, it's definitely possible to envision a more than interesting future.