Is Attention All AGI Needs?
I keep coming back to this question: is the transformer enough? If we keep scaling it, does it eventually become something that actually reasons, something you could honestly call general intelligence?
My short answer is probably not on its own. But I also don't think the architecture is the main thing holding us back right now.
It's a simpler thing than people think
A transformer is attention so every token can look at every other token, an MLP, and a lot of those stacked on top of each other. That's most of it. If you've ever written one from scratch, the model code ends up being one of the shortest files in the repo. Everything around it, the data, the tokenizer, the training loop, is where the real work is.
To me that says a lot. The architecture is a very good general-purpose learner that happens to scale really well on GPUs. The intelligence isn't sitting in the block itself. It comes from what you train it on and what you ask it to predict.
The case against it
The criticisms are real, and I don't want to brush them off.
A transformer spends the same amount of compute on every token, whether the problem is trivial or really hard. It has no memory past its context window. It stops learning the moment training ends. It only ever sees the world through text or pixels, never by doing something and watching what happens. And next-token prediction rewards sounding right, which is not the same thing as being right.
For a while these felt like dealbreakers to me.
Then reasoning models showed up
Letting a model think for thousands of tokens before it answers, and training that with RL on problems where you can check the answer, fixed a lot of the fixed-compute problem without changing the architecture at all. The chain of thought turns into a scratchpad, and the model gets to spend more compute on harder questions. Same transformer, different objective, completely different behavior.
That pushed me toward the view that it's mostly the training, not the architecture. A lot of things people said transformers could never do turned out to be things we just hadn't trained them to do yet.
Where I think it still breaks
Still, I don't think scaling plus RL gets us all the way there.
The biggest gap for me is continual learning. We don't retrain people from scratch every few months. We pick things up from a handful of examples and we don't forget everything else when we do. Models can't really do that yet, and stuffing more into the context window is a workaround, not an answer.
The second is grounding. I care a lot about world models and robotics, and I don't believe you get solid common sense about physics and cause and effect just by reading about it. At some point a system has to act, predict what happens, be wrong, and update. That's a very different loop from how LLMs are trained today.
The third is efficiency. A kid learns language from a tiny fraction of the data we throw at these models. A gap that big means something about how learning works is still missing, and I don't think more parameters is what fills it.
So is it enough?
My guess is that whatever gets us to AGI will still have something like attention inside it. Attention is too useful to throw away. But it'll be one piece of a bigger system, one that has real long-term memory, keeps learning after it's deployed, learns from acting and not only from text, and can decide for itself how long to think.
Whether we still call that a transformer is mostly a naming question. The block might survive. The recipe around it won't look like today's.
That's honestly the part I find most interesting. Not transformers versus some new architecture, but figuring out what the rest of the system needs to be. It's a big reason I like building models from scratch, so I can see for myself which parts actually matter and which ones we've just gotten used to.
This is where I'm at right now. Ask me again in a year and some of it will probably be wrong, which is kind of the fun part.