Parascandolo's 2020 MIT interview presentation proposed three key ideas: open-ended reasoning (investing more compute for harder problems), language as reasoning carrier, and system self-modification—concepts now central to modern reasoning AI. The researcher, trained under Bernhard Schölkopf at ETH Zurich with internships at Google X and DeepMind, joined OpenAI's reinforcement learning team and later led work on the o1 and o3 projects.
MIT's "Nonsense" Rejected 2020 Presentation Now Underpins OpenAI's o1 and o3
The model must learn how to use additional compute, not just have it.
Why did MIT's professors dismiss this as nonsense in 2020? What was the intellectual barrier?
The field was still operating under the assumption that you built a model once, fixed its architecture, and then ran it. The idea that you'd let a model think for variable amounts of time—that you'd scale compute with problem difficulty—didn't fit the paradigm. Transformers were working. Why complicate things?
But Parascandolo had data showing RNNs couldn't just loop indefinitely and get better. He wasn't being naive.
Exactly. He was identifying a real problem: more compute alone isn't enough. The model has to learn how to use it. That's a harder problem than it sounds, and it required new training techniques and new ways of thinking about what a model is doing when it reasons.
The language-as-reasoning idea seems almost obvious in retrospect. Chain of Thought became huge.
It does seem obvious now. But in 2020, the dominant view was that language models were pattern-matching machines. The idea that you could use language as a reasoning substrate—that the model could literally think in words—was less obvious. Parascandolo was saying language isn't just output; it's the medium of thought itself.
What about the third idea, the one about systems modifying their own weights? That sounds speculative even now.
It is. But it's also the most interesting. He was proposing that learning becomes part of the action space. The system doesn't just learn during training; it learns during deployment. It can generate its own training data, construct scenarios, retry failed paths. That's closer to how humans actually improve.
How much of o1 and o3 is actually built on these ideas versus just convergent evolution?
That's hard to say from the outside. But Parascandolo was there for all of it. He moved from reinforcement learning to the algorithm team to the Strawberry team—that's the o1 project. He shaped the direction. Whether o1 would exist without his 2020 presentation is unknowable. But the conceptual framework he laid out is definitely there.
What does it say about academic gatekeeping that this got dismissed?
It says that novel ideas often look wrong before they're proven right. The professors weren't stupid. They were just operating within existing constraints. Parascandolo had the advantage of thinking beyond those constraints—and then the resources and time to actually build what he imagined.
Der Puls
- 2020 MIT presentation dismissed as 'nonsense' by faculty committee
- Proposed three core ideas: compute-time scaling, language-based reasoning, system self-modification
- Parascandolo joined OpenAI in 2021, worked on o1 and o3 projects
- PhD under Bernhard Schölkopf at ETH Zurich; internships at Google X and DeepMind
Parascandolo's 2020 MIT interview presentation proposed three key ideas: open-ended reasoning (investing more compute for harder problems), language as reasoning carrier, and system self-modification—concepts now central to modern reasoning AI. The researcher, trained under Bernhard Schölkopf at ETH Zurich with internships at Google X and DeepMind, joined OpenAI's reinforcement learning team and later led work on the o1 and o3 projects.
A 2020 presentation by OpenAI researcher Giambattista Parascandolo, dismissed as "nonsense" by MIT professors, outlined core concepts now fundamental to reasoning models like o1 and o3, including compute-time scaling and language-based reasoning.
Five years ago, a researcher named Giambattista Parascandolo walked into MIT for a faculty interview and presented his ideas about how to build reasoning into neural networks. The professors on the committee dismissed the work as nonsense. Today, those same ideas form the conceptual backbone of OpenAI's o1 and o3 reasoning models—systems that represent some of the most advanced AI development happening anywhere.
Parascandolo has since made the incident public, posting about it on his personal homepage along with slides from that 2020 presentation. What he was proposing then sounds almost prescient now. The core insight was simple but radical for its time: a model should be able to spend more computational effort on harder problems. A difficult question might require many steps of reasoning; an easy one might need just one. The amount of thinking should scale with the problem's complexity, not be fixed by the model's architecture.
At the time, this ran against the grain of how neural networks were built. Standard Transformers had fixed depth—every token went through the same number of layers regardless of whether the task was trivial or demanding. Recurrent neural networks seemed like a better fit, since they could loop and theoretically think for as long as needed. But Parascandolo's data showed a problem: RNNs typically peaked in accuracy at the number of reasoning steps they'd seen during training. Add more loops beyond that, and performance actually dropped. The implication was clear: simply giving a model more compute wasn't enough. The model had to learn how to use that compute effectively.
The second pillar of his thinking was language itself. Parascandolo argued that language could serve as the medium through which a neural network reasons. He used the game Montezuma's Revenge as an example—a reinforcement learning agent trained from scratch needs to explore countless state-action combinations, but what it really lacks is judgment about which behaviors make sense. GPT had already absorbed vast amounts of world knowledge from text. Language could help a model describe its environment, understand goals, break tasks into pieces, and generate high-level plans. It could narrow the search space dramatically. This line of thinking maps directly onto what we now call Chain of Thought prompting and agent workflows.
The third direction was perhaps the most speculative: systems that could reset tasks, return to earlier states, construct counterfactual scenarios, and even modify their own weights and activations. In essence, Parascandolo was proposing that learning itself become part of what the system could do—deliberate practice, generating training scenarios on the fly, transferring knowledge, relearning from failures. It was a vision of AI systems that could introspect and improve themselves.
Parascandolo's background explains why he was thinking this way. He did his doctorate at ETH Zurich under Bernhard Schölkopf, one of the field's leading theorists on generalization and out-of-distribution learning. During his PhD, he interned at Google X on large-scale simulator design and at DeepMind on Monte Carlo Tree Search. When he graduated in September 2021, he joined OpenAI directly, moving through the reinforcement learning team, then the algorithm team, and eventually to what was internally called the Strawberry team—the codename for the o1 project. He's since worked on GPT-4 and the foundational research behind both o1 and o3.
What makes this story worth attention isn't just that Parascandolo was right and the MIT professors were wrong. It's that the gap between dismissal and vindication was five years, and that vindication required building entire new systems, training them at scale, and proving the concepts worked in practice. The ideas weren't obviously wrong in 2020—they were just ahead of what the field knew how to do. A blog post he wrote in 2021 offers another window into his thinking. He pushed back against the claim that neural networks need massive data because they're fundamentally different from human brains. His argument: humans don't learn to recognize dogs from scratch either. Evolution has already done the heavy lifting, encoding ancestral experience into our brains' structure. Large-scale pretraining of neural networks is analogous to that evolutionary process. Fine-tuning and in-context learning are like individual learning. From this view, continuing to expand data and compute might still yield significant gains—not because the mechanisms are identical to biology, but because the principle holds.
Parascandolo remains relatively quiet. His last post on X was in November 2024, when he was recruiting research engineers and software engineers for o1. Some of the technical details on his homepage are still redacted with black blocks. But the arc is clear: a researcher with deep roots in generalization and planning theory proposed a framework for reasoning that seemed implausible to established academics, then spent the next five years helping build systems that made it real.
Bemerkenswerte Zitate
The model can invest more time and computing power to continuously revise its answers. The more difficult the problem is, the more steps the model should think through.— Giambattista Parascandolo, 2020 MIT presentation