Beyond the Turing Test: The New Age of Narrow AI Mastery
A few people asked me to clarify what “narrow AI” actually means, so let us start there.
Narrow AI refers to systems that are highly skilled at specific tasks without possessing general intelligence or broader awareness. They do not “think” like humans. They cannot reflect, feel or draw meaning from lived experience—yet. But they can outperform us in tightly defined domains such as transcribing speech, writing summaries or answering questions.
Over the past few years, they have become very good at those tasks, very quickly.
The chart below, based on Mensa-style quiz data from TrackingAI.org, illustrates the pace of that shift. Historical snapshots put language models around IQ-equivalent scores of 40 on these tests in 2022. By 2023 they were scoring around the human average of 100. By 2024–25, models such as Claude 3 Opus and Gemini 2.5 Pro reached benchmark-specific equivalents of 120–150+, placing them in the top few per cent of human test performance.
How did we get here so fast?
In 2022, models such as GPT-3 were good at generating fluent text but poor at reasoning. They failed logic puzzles, misunderstood subtle instructions and made basic errors. Their IQ-test equivalents were closer to 40–70, far below the average human. They were impressive in one sense, but they did not work reliably. I was fortunate to co-author an award-winning ACL paper showing that GPT-3 did not always work as expected.
In 2023, things shifted dramatically. GPT-4 and other new models scored in the upper quartile on exams such as the SAT, GRE and LSAT. For the first time, language models could pass professional certifications, summarise dense legal texts and tutor students with consistent accuracy.
Then came 2024–25 and, with it, Claude 3 Opus, Gemini 1.5 and 2.5 Pro, and GPT-4o with vision. These models did not merely match average humans. They began to outperform almost all of us in tasks involving verbal reasoning, pattern recognition and some structured logic—especially when speed is taken into account.
But IQ is not intelligence
These “IQ scores” are useful benchmarks, but they do not mean machines are necessarily smarter than us. They measure narrow, test-based forms of reasoning, not the wider scope of human thought. I do think these models encode some kind of “world model”, but it is extraordinarily impoverished. Whether and in what sense they understand the world remains an open scientific debate.
This is where the term “narrow AI” matters most. These models have a limited understanding of the world, but what they can do with that narrow skill set is still remarkable.
And now, the Turing Test?
In March 2025, researchers ran a modern version of the Turing Test—Alan Turing’s benchmark in which a machine attempts to convince a human judge that it is a person.
A system based on GPT-4.5 was judged to be the human participant in 73% of conversations under the study’s conditions.
We are entering an era in which machine language fluency can be indistinguishable from human fluency in many settings. This has real implications for trust, information, education and the future of work.
So where are we now?
Several things are clear:
- Narrow AI now outperforms most humans in specific language tasks.
- IQ-equivalent benchmark performance has risen from roughly 40 to 150 in a few years.
- We are entering an age where “passing the test” is not the limit; it is the baseline.
The question is no longer whether AI can beat us at specific tasks. It can, and it does. The challenge is what we do with that capability and how we redesign our responsibilities around it.
This is not a retreat from human intelligence. It is a rethinking of intelligence in a world where we are no longer the only ones who can wield it.
As we cross from 40 to 100 to 150+ in a few short years, it is worth remembering: the models are rising—the real question is how we rise with them.
References
- Jones, C. R. and Bergen, B. K. (2025). “Large Language Models Pass the Turing Test”.
- Lu, Y. et al. (2022). “Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity”. Proceedings of ACL.