As AI evolves, experts debate the relevance of traditional IQ metrics in measuring machine intelligence. New benchmarks may be needed to respect AI’s unique capabilities.

In a world where artificial intelligence is reshaping the landscape of possibility, we find ourselves at a curious crossroads.
The question that hangs in the air: How do we measure the intelligence of machines?
While the concept of IQ has long been used as a yardstick for human intelligence, its application to AI is ruffling feathers across the tech world.
Recently, Sam Altman, CEO of OpenAI, made waves with his comments suggesting that AI’s “IQ” improves by roughly one standard deviation each year.
But this casual comparison stirred a hornet’s nest of debate.
Can the intricacies of human intelligence really be mirrored in machines, and is IQ the right tool for the job?
Sandra Wachter, a voice of reason from Oxford, reminds us of the apples-and-oranges nature of this comparison.
While IQ tests can somewhat capture human logic and abstract reasoning, they fall short in evaluating practical intelligence—the very essence of getting things done.
When applied to AI, these tests become a reflection of their own limitations rather than a true measure of a model’s capabilities.
The origins of IQ tests, steeped in the murky waters of eugenics, further complicate their standing as a fair measure.
As Os Keyes from the University of Washington points out, AI models, with their near-limitless memory and patience, can essentially “game” these tests.
This is akin to a marathon runner measuring their speed against a cheetah; it’s an unfair race from the start.
Moreover, AI models often train on vast swathes of public data, which may include countless IQ test examples.
This means they’re not so much solving the test as they are recalling pre-learned patterns, as Mike Cook from King’s College London elucidates.
Unlike human brains, which juggle distractions and limitations, AI operates in a realm of perfect recall and clarity.
The narrative here isn’t just about the inadequacy of IQ tests for AI but about the broader need for new benchmarks.
Heidy Khlaaf from the AI Now Institute suggests we shift our perspective.
The historical avoidance of comparing machine capabilities to human abilities is telling.
Machines have always excelled in ways humans cannot—solving problems with speed and precision beyond our reach.
This is not to say machines “think” like us.
Rather, they operate on a fundamentally different plane.
The challenge lies in developing metrics that respect these differences, evaluating AI on its own terms.
As we stand on the brink of what AI could become, let us remember that intelligence is not a monolith.
It is a tapestry of diverse capabilities, each deserving of its own lens.
In the end, the journey of understanding AI’s “intelligence” is as much about redefining our own notions of intelligence as it is about the machines themselves.
The future beckons, and with it, the promise of a new era of discovery—one where we learn to see beyond the constraints of traditional measures and embrace a broader, more nuanced understanding of what it means to be intelligent.