MusicLM: Artificial Intelligence to Generate Music from Text
Artificial intelligence for generating music

In mid-December we talked about Riffusion, a Stable Diffusion variant that lets us create music with artificial intelligence from simple text. Time is on the side of algorithms, and with each new project we get more precise results. Today it's Google Research's turn, which just unveiled MusicLM. Besides creating audio at 24 kHz, this model supports special conditions like generating long melodies and a story mode to prepare sequences.

The viral explosion of ChatGPT triggered a code red at Google. The Mountain View giant uses artificial intelligence at various levels, but everything seems to indicate that a wave of new projects is approaching as a direct response to OpenAI's chatbot. In other words, Google needs to show the public its cards a bit more, and one of them is MusicLM. This work by Google Research offers something very interesting: Generate music with artificial intelligence using a simple description.

What Does AI-Generated Music Sound Like?

The demo page does not have an active MusicLM model, but it is full of examples accompanied by their respective prompts. MusicLM's training is based on "a large dataset of unlabeled music", and on another dataset called MusicCaps, with a total of 5,521 music-text combinations. The MusicCaps descriptions were created by humans, and the associated audio comes from AudioSet, a collection with more than two million audio clips (ten seconds long), extracted from YouTube.

MusicLM: Artificial Intelligence to Generate Music from Text
MusicLM can take inspiration from paintings

"The main soundtrack of an arcade game. It is fast-paced and upbeat, with a catchy electric guitar riff. The music is repetitive and easy to remember, but with unexpected sounds, like cymbal crashes or drum rolls."

The MusicLM results are divided into several categories. The first of them is "Rich Captions", with 30-second samples based on a brief description. "Long Generation" shows the model's potential to create complete songs lasting five minutes. "Story Mode" turns descriptions into sequences with defined intervals, "Text and Melody Conditioning" combines the prompt text with a reference melody, and "Painting Caption Conditioning" generates audio inspired by a painting or image.

MusicLM: Artificial Intelligence to Generate Music from Text
Story Mode turns the prompt into a sequence to follow

In addition, MusicLM can reproduce specific instruments and genres, levels of musical experience (a novice pianist or a master violinist), places, eras, and more. However, all we have so far are its official examples. Google Research confirmed that "there are no plans" to share models for now, and they have many challenges ahead (hypothetical copyright issues, cultural bias, etc.).

**Official site:** Click here.