It’s a matter of time, we all know it. AI video generation is in a basic, almost primitive state... but the same was said about image generation a year ago, and today we can't stop using its algorithms. Recently, the company Runway shared a teaser on YouTube and Twitter, anticipating the potential of its Text to Video tool. The demo looks great... but what can we expect in real life?

The Next Phase: AI Video Generation with Text
Video generation

AI image generation can lead to enormous differences depending on the algorithm used, but in general they all work the same way: First we describe what we want (a portrait, a landscape, etc.), and then we add secondary elements that serve as influence or inspiration, such as an artist's name, a specific genre, a painting style, and even the name of a camera.

Now, imagine that when creating video. We enter the editor and ask for a city with a lot of traffic. Then we request a taxi embedded against a traffic light, and finally the reason for the crash: A giant dinosaur running by. Of course, a video editing expert can do that today with the right budget, but Runway suggests that any user could get similar results in the future through its new tool, Text to Video.

Runway Text to Video: Creating Videos with Artificial Intelligence and Text

A city street, a cinematic filter, a light pole that disappears by magic. A beautiful garden, multiple images, dynamic text injection. A 'green' character, a blurred background in real time, and a touch of Robert Capa as seasoning. The teaser is short but promising, and it also helps us define Runway's vision. Instead of training an independent freely accessible algorithm, the company seems to aim at developing a more advanced editor with the integrated text-to-video function.

If you want to be part of the early access, all you have to do is visit the official page and fill out a form. However, all this raises a question: How are they going to train such a model? Currently there is nothing similar to LAION-5B for video, and the level of processing required is immense. We will be waiting...

Sources: Runway, Ars Technica