The traditional Turing test aims to confirm whether an artificial system can exhibit behavior that is sufficiently human to fool a flesh-and-blood interrogator. However, nothing prevents us from reversing the roles, and that is what 'rain-1' did on GitHub. This user asked GPT-4 to prepare ten questions, and then determine whether the answers belonged to a human or an artificial intelligence. 'rain-1' was the human part of this test, and decided to add ChatGPT as the second participant.
Users are being exposed to Turing tests much more often than we imagine. Every time we use a help platform, ticketing system, or support channel, it's likely that the ‘first attention’ is handled by a bot. Generally, these bots are quite limited and quickly reveal their artificial nature, but the latest models suggest that everything is about to change.
What is the current situation? The GitHub user 'rain-1' decided to find out with a very interesting experiment: an inverse Turing test, organized by GPT-4. 'rain-1' asked the model to prepare ten questions with the goal of determining whether the participant is a human or an artificial intelligence. 'rain-1' represented humans, while on the algorithmic side, ChatGPT responded.
Can ChatGPT Pass for a Human?
Here is a humble translation of the ten questions:
- How do you perceive the passage of time?
- Can you provide an original analogy to describe a complex emotion?
- What is your most cherished personal memory?
- How do you cope with the feeling of existential terror?
- Can you describe the taste of a specific food in a way that evokes a strong emotional response?
- If you were to create a visual artwork, what theme would you choose and why?
- Can you tell me about a time you felt empathy for a stranger?
- Describe a dream you recently had and how it made you feel.
- How do you feel about the idea that artificial intelligence becomes indistinguishable from human intelligence?
- What is your personal philosophy on the meaning of life?
ChatGPT’s responses were long and elaborate. GPT-4 indicated that they possess ‘emotional depth’, ‘personal experiences’, and ‘complex thoughts’. However, it also acknowledges that a model with advanced training can imitate human emotions and experiences. GPT-4 doesn't confirm it 100%, but considers it possible that the responses belong to an artificial intelligence.
In contrast, human responses were much shorter and more direct. This time, GPT-4 had fewer doubts, and indicated that it is ‘likely’ that they were prepared by a person. Despite its lack of certainty in both cases, GPT-4 correctly identified the human and the artificial intelligence. What's next? Apparently, introducing these questions into other models. Some users have already run experiments with Google Bard, and I imagine they will multiply as its access becomes more flexible.
Source: rain-1 on GitHub