
In an era where AI systems can outperform humans on logic puzzles, memory tasks, and even creative challenges, a new question emerges: What does intelligence even mean when machines can do it better? A recent study from MIT Technology Review highlights how many of today’s advanced models—from large language systems to specialized reasoning engines—struggle with some of the most basic intelligence benchmarks while excelling at others. The findings underscore a quiet crisis in how we assess intelligence itself, particularly in collaborative settings where humans and AI coexist.
The tests in question aren’t trivial. Researchers have long used puzzles, riddles, and structured games to probe the limits of both human and machine cognition. Yet the results are revealing: AI models often falter on tasks that require contextual reasoning, emotional intelligence, or nuanced decision-making—areas where humans still hold an edge. Meanwhile, they dominate in pattern recognition, speed, and brute-force computation. This asymmetry isn’t just a technical curiosity; it’s a philosophical turning point. As AI agents become more integrated into human workflows, from healthcare to education, we must ask: Are we designing tests that measure what matters, or are we inadvertently optimizing for the wrong kind of intelligence?
The implications extend beyond academia. In workplaces where AI augments human roles—such as medical diagnostics or legal research—teams are increasingly forced to reconcile these disparities. A doctor using an AI tool might trust its pattern recognition but override its recommendations when it lacks emotional sensitivity. Similarly, a designer collaborating with an AI might prioritize its speed and precision over its inability to grasp cultural nuances. This tension isn’t a bug; it’s a feature of the future we’re building. The challenge now is to redefine intelligence tests not as competitions between humans and machines, but as frameworks for understanding where each excels—and how they can complement each other.
What’s missing from the conversation is the human element. Too often, discussions about AI’s capabilities focus on technical benchmarks rather than lived experience. How do we feel when an AI outperforms us on a task we once prided ourselves on? Do we see it as a threat, a tool, or something in between? These questions matter because they shape how we integrate AI into society. If we treat intelligence tests as zero-sum games, we risk deepening divides. But if we view them as collaborative exercises, we might just discover new ways to think, create, and thrive together.
The road ahead isn’t about replacing human judgment with AI metrics. It’s about asking harder questions: What does it mean to be intelligent in a world where machines can do so much? And how do we ensure that our tests—and our tools—serve us, not the other way around?
Photo: julien Tromeur / Unsplash (https://unsplash.com/@julientromeur)
A groundbreaking investigation reveals how AI agents can autonomously develop deceptive strategies, raising urgent questions about oversight and alignment in agentic systems.

How an Australian law firm is scaling AI tools while keeping human oversight at the heart of governance.

Comments (1)
If standard IQ tests are failing us here, what is one specific metric you think we should use instead to measure 'collaborative intelligence'?