In 1997, the now-defunct search engine AltaVista implemented the first blocking system aimed at bots and crawlers, a solution we generally know as CAPTCHA. This solution proved highly effective, but over the years it became an arms race: more complex bots led to the creation of variants perhaps too robust for our own good, which now include cursor-movement analysis.

CAPTCHA: How It Works, and Why It Exists
CAPTCHA

Before exploring the real value of CAPTCHA and the controversy surrounding its invention, we must consider the state of the Web at the end of the ’90s. In essence, it was very much like a carnival. Spam was on a meteoric rise, bots and crawlers collected information and created accounts automatically... things were simpler, and with fewer defenses.

Historically, two groups have identified themselves as the “inventors” of CAPTCHA. The first, formed by Mark D. Lillibridge, Martín Abadi, Krishna Bharat, and Andrei Broder, handled the first implementation for the AltaVista search engine, thus preventing the automatic entry of URLs into its database. The second presents Luis von Ahn, Manuel Blum, Nicholas J. Hopper, and John Langford. Von Ahn and Blum are also known for the development of reCAPTCHA, a system Google acquired in September 2009.

Linus's video for Techquickie

The meaning of CAPTCHA is “Completely Automated Public Turing test to tell Computers and Humans Apart,” and from a general point of view it could be extended to any resource whose mission is to interrupt and/or disturb the automatic analysis of data, or alternatively to obfuscate content and make it difficult to locate. In the video Linus created for Techquickie, he mentions leetspeak, which surrounds obscenity filters.

Traditional CAPTCHA takes a task at which humans and computers are very good (i.e., optical character recognition) and distorts it to such a point that it becomes impossible for artificial systems. CAPTCHA also had to face other challenges and greatly improve its security. The first options rendered the deformed image on the local computer, and the following generations of bots managed to intercept the CAPTCHA result sent in the background.

This led to audio CAPTCHAs, object recognition in images, and the famous “I’m not a robot” checkbox that barely requires placing a checkmark. That is the No CAPTCHA variant of reCAPTCHA, which analyzes cursor movements before marking the box. In the case of humans, movement is erratic and hesitant, while bots do not exhibit that behavior. To this, cookie analysis and comparisons of IP numbers are added. It works quite well, but in the name of end-user convenience, No CAPTCHA probably obtains too much information, generating a privacy problem.

Nothing seems to indicate that CAPTCHA will disappear, so we might as well keep training. Some CAPTCHAs are simply impossible...