Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
Text to speech

The algorithmic explosion has allowed us to create images and music, but it also showed some strength in the transformation of speech to text. If someone is looking for a subtitle or a transcription, there are many free services available, but the question is, how does it work in the other direction? I'm referring to text-to-speech, the classic text-to-speech conversion that a large number of users need frequently. Today we're going to explore some online platforms, study their free modes closely, and as far as possible, compare results.

Text to Speech: Better with Each Generation

I still remember when the first version of Dragon NaturallySpeaking hit the market. We had to fight for hours to train the system, and with a bit of luck, it would end up recognizing 25 percent of the words. These days, speaking to a device and transmitting verbal commands is so simple that nobody stops to think about the complex evolutionary process that brought us here.

Speech synthesis is even older (we can think of the Bell Labs vocoder, or some IBM computer singing Daisy Bell in the '60s), but with the emergence of new models based on artificial intelligence, it's not crazy to say that it's at its best. In fact, text-to-speech conversion is closer than ever, and today we're going to explore the free offerings of some platforms specially designed for this task.

How to Convert Text to Speech: ElevenLabs

Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
The basic demo is good, but there are more advanced commands inside

The free mode of ElevenLabs limits its processing to ten thousand characters per month and does not authorize commercial use, but that should be more than enough for any personal project. The list of available languages extends to 28, while the enabled voices in the demo (maximum of 333 characters) are 29.

ElevenLabs demo

This demo is not only impressive on its own, but ElevenLabs also allows us to download a copy of the result in MP3 format. To enter the platform, the simplest way is to use Google credentials, and inside we find advanced parameters such as alternative models, stability, and clarity. Now, ElevenLabs is not perfect (some voices convert 25 to 'twenty-five' instead of 'veinticinco'), but it offers an excellent starting point for all kinds of users.

Speechify

Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
This is Speechify's "basic mode"...

Speechify is a different creature than the rest. From a certain point of view, it divides into two parts: One focused on traditional text-to-speech with nine voices in Spanish, and the full "webapp", where we must input text such as a link to a web page, a new document, a local file, a scanned physical document, or a copy saved in the cloud. In other words, Speechify is a productivity tool that will help us process texts faster.

Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
... but the webapp isn't bad.
Speechify demo

The Spanish language has a total of twelve voices in the webapp, and I personally recommend trying them all. Quality differences can be very significant even among similar voices, and speed control is essential to optimize results. Speechify does not enable audio download in its free mode, but nothing stops us from recording it with a compatible tool in the background.

Bark

Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
You need to answer a couple of questions on Discord before starting...

Bark is a creation of Suno AI, a platform much better known for its virtues in generating music with artificial intelligence through the Chirp tool. Its text-to-speech technology is open source, which means any Python rider can go and explore the code on GitHub as well as explore demos, samples, and other technical details.

Text to Speech: Convert Entire Texts into Audio with Artificial Intelligence
The commands aren't entirely intuitive, but they're learned relatively quickly
Bark demo, random voice

The simplest alternative is obviously to head to its Discord server, enter the Bark beta channel, and start generating voices. The commands are divided into /bark to activate the engine, followed by prompt (the text we want to process), and voice, where we can roll the dice using the random option, or choose a voice from the list. In less than a minute, Bark will share the conversation, available in MP3 and MP4.

In Summary

I hope this selection and its samples serve to give you a solid idea about text-to-speech conversion using artificial intelligence, and what the most important limitations are. Beyond the inevitable restrictions in free profiles, I think the conditions are definitely appropriate to cover most of our needs, without spending a single cent.