On February 14, YouTube turned 20 years online. According to the most recent estimates, it has surpassed the 2.5 billion users mark, a staggering number by any measure. However, Google is very reluctant to share statistics. If we want to know precise details like the real number of published videos, the most common topics, or the average views, the official answer is… silence. Fortunately, a team from the University of Massachusetts Amherst accepted the challenge of obtaining firmer data, and their solution is (believe it or not) to randomly load content.
On February 11, YouTube CEO Neal Mohan published an article on the official blog of the platform, highlighting the extraordinary growth of YouTube on American screens. In fact, the classic television surpassed mobile devices for the first time in that country when considering "watch time" across all YouTube services. Personally, I have witnessed this: many people with Smart TVs set up their playlists on YouTube and play them there, completely ignoring traditional cable, digital television, and FAST options.
Now, if you want to know other statistics about YouTube… good luck. Apparently, Google doesn't think it's a good idea to reveal the true dimension of YouTube, and is very comfortable managing the platform as if it were a black box. But there is another way to obtain more solid data… and it's one video at a time. That brings us to the University of Massachusetts Amherst, where Ethan Zuckerman (director of the Initiative for Public Digital Infrastructure) and his team created a program that generates random YouTube links… billions at a time.
Analyzing the "Deep YouTube"
Zuckerman explains that the program essentially works similarly to a scraper, although other sources have associated it with the old term of wardialing. As we already know well, each video on YouTube has a unique string of eleven characters. For example, "dQw4w9WgXcQ" doesn't need any introduction, but Zuckerman's program generates eleven random characters and checks if a video is available. The problem is that the number of possible combinations amounts to 2^64 (a number we know well for other reasons), so they had to settle for a smaller dataset. In their original study, to obtain a sample size of 10,016 unique videos, they processed about 1.9 billion erroneous links for each correct one.
So, what can we calculate from that? Although the data goes up to June 2024 (with a sample size of 25,953), the estimated number of videos available on the platform amounts to 14.83 billion, a 60% jump compared to 2022. However, this clashes with other important details:
- Only 0.21% has some kind of monetization
- Less than 4% invites the audience to subscribe, comment, or leave a like
- Only 38% received some kind of editing
- 18% has high-quality audio
- More than 40% is solely music, without dialogue
- 16% of videos is composed of static images
- The average number of views is 41, with a duration of 64 seconds
- 74% of videos have no comments, and 89% don't register a single like
- 4% of videos didn't have a single view
In short, YouTube has stopped being that space that invited us to "broadcast ourselves". When Neil Mohan said "YouTube is the new television", he wasn't exaggerating. The "Deep YouTube" is essentially a desert of short, erratic, and unedited videos that nobody watches. If you wonder why the algorithm keeps favoring the Mr. Beast types of our lives… I think here is the answer.
Access the original study: click here.
Official site: click here.