The past few years have been an exciting time in AI. Tools like GPT-4.5 and Gemini Ultra have redefined what’s possible with language models. They answer questions, summarize reports, write code, and even generate creative content at a level that felt impossible just a few years ago.
But as impressive as these models are, there’s a growing gap between expectations and reality when it comes to Artificial General Intelligence (AGI). The dream of machines that can reason, learn, and adapt like humans is still just that – a dream. And unless something fundamentally new happens in the AI space, LLMs aren’t going to get us there.
Bigger Models, Same Limitations
Let’s start with the obvious. Most progress in LLMs comes from making models bigger and feeding them more data. That’s led to longer context windows, more fluent answers, and better language predictions. But the core method hasn’t changed. It’s still all about next-word prediction based on patterns in training data.
The jump from GPT-3 to GPT-4.5 shows this. More memory, better coherence, but no fundamental shift in how the model “thinks” – because it doesn’t actually think.
Apple’s Study: Proof LLMs Aren’t Reasoning
If anyone needed more proof that LLMs fall short of real reasoning, Apple just delivered it.
In their 2025 study, “The Illusion of Thinking,” Apple’s researchers tested top models like GPT-4, Claude 3.7, and DeepSeek-R1 on classic logic puzzles – Tower of Hanoi, river-crossing problems, and others.
The results? These models handled easy tasks. But when complexity increased, their performance dropped to almost zero. Worse, when the questions got harder, the models didn’t produce more reasoning steps. They actually gave up sooner.
Apple’s conclusion was blunt: today’s LLMs aren’t reasoning – they’re mimicking. They’re pulling statistically likely answers from training data without real problem-solving ability. Even small distractions in a question caused up to a 65% performance drop, showing how easily these models get thrown off track.
For anyone watching AI’s long-term trajectory, this is a wake-up call. Scaling isn’t solving the reasoning gap.
Why Diffusion Models Matter Here
If you want to see what a real breakthrough looks like, just look at the image generation space.
When Stable Diffusion hit in 2022, it opened new creative possibilities for millions of users. But the real leap came with Midjourney V7 in 2025. Midjourney didn’t just scale up. They rethought how models handle style, nuance, and user intent. The result wasn’t just sharper images – it was a whole new level of realism and artistic control.
| Model | Approach | Key Leap | Outcome |
|---|---|---|---|
| Stable Diffusion | Diffusion-based | Democratized image synthesis | High-quality, accessible image creation |
| Midjourney V7 | Novel architecture | Breakthrough in realism and intent capture | Photorealistic, nuanced output |
The lesson: Real progress happens when you change the approach – not just when you make something bigger.
What AGI Actually Needs
If AGI is the goal, future models need to do more than generate text that sounds right. They’ll need to:
- Understand cause and effect
- Learn and adapt in real time
- Integrate multiple forms of input (text, images, sound, action)
- Reason through complex, multi-step problems
- Separate signal from noise in messy, real-world data
Right now, LLMs don’t check any of those boxes.
What Business Leaders Should Take Away
For leaders making technology decisions, the message is simple: Enjoy the productivity gains LLMs bring today, but don’t confuse scale with intelligence. If your roadmap includes complex decision-making AI or autonomous reasoning tools, you’ll need to watch for true breakthroughs – something more like what we’ve seen in the image generation world.
Apple’s study makes it clear: Without a new kind of thinking, LLMs won’t close the gap to AGI. The field needs new ideas, new architectures, and new ways to teach models how to reason – not just how to predict.
Bottom line: Bigger isn’t always smarter. The next wave of AI progress will come from invention, not repetition.