Neel Nanda
Runs the mechanistic interpretability team at Google DeepMind; worked on interpretability at Anthropic under Chris Olah before that, and teaches the field to newcomers.
Neel Nanda did not write this page. What is this?
It collects the places they publish and what they have said there, each linked to the source. They have no account here. Is this you? Claim it, correct it, or ask us to remove it from ppll.
Where they publish
Beliefs
Korrents What they believe 17 beliefs — each backed by an exact quote.
Each is a — compiled by korrents.com, not by them: the one-line wordings are korrents', the quotes are theirs.
Recent
Complete reverse-engineering of neural networks is a less promising research direction than studying model biology and applying interpretability usefully
I'm now fairly pessimistic about ambitious interpretability (i.e. complete reverse-engineering), and I'm excited about model biology (studying qualitative high-level properties of models) and applied interpretability (rigorously doing useful things with interp).
MATS Applications Open (Due Aug 29) Said 19 Aug 2025
Sparse autoencoders are a useful interpretability tool but are often used wastefully when a simpler, more obvious method would work as well or better
I'm more agnostic about the best techniques, things like sparse autoencoders are a useful tool, but easy to waste effort using when a simpler method is sufficient or better - start by doing the obvious thing!
MATS Applications Open (Due Aug 29) Said 19 Aug 2025
Research taste, the harder-to-learn conceptual skills behind good research strategy, takes a long time to develop but very little time to apply once acquired
My model is that research requires a mix of skills. The day-to-day coding and execution is crucial. But there's also a set of harder-to-learn conceptual skills, collectively called research taste. These skills take a long time to gain because they have poor feedback loops, but they take very little time to use.
MATS Applications Open (Due Aug 29) Said 19 Aug 2025
Show 14 more
The person receiving advice usually knows more about their own situation than the person giving it does.
No matter how much I know about a domain, the other person will always know far more about their own situation, context, beliefs, skills, preferences, etc than I do.
Post 51: Socratic Persuasion: Giving Opinionated Yet Truth-Seeking Advice Said 26 May 2025
People engage more with an argument when they generate its key steps themselves than when the same argument is imposed on them.
It's a lot easier for someone to engage with an argument if they generated the key steps themselves by answering my questions - if imposed by me, it sparks contrarianism and defensiveness
Post 51: Socratic Persuasion: Giving Opinionated Yet Truth-Seeking Advice Said 26 May 2025
A strong track record in a technical field is only weak evidence that someone has good strategic judgment about AGI.
Having a good research track record is some evidence of good big-picture takes about AGI, but it's weak evidence.
Post 50: Good Research Takes are Not Sufficient for Good Strategic Takes Said 22 Mar 2025
People's opinions are shaped more by who they talk to most than by what is actually true.
Though note that people's opinions are often substantially reflections of the people they speak to most, rather than what's actually true.
Post 50: Good Research Takes are Not Sufficient for Good Strategic Takes Said 22 Mar 2025
The best way to maximize impact is often through high-risk, high-reward strategies.
often the best way to maximise your impact is by pursuing high-risk, high-reward strategies
Post 49: Things That Make Me Enjoy Giving Career Advice Said 17 Jun 2022
Careers in AI and effective altruism are highly competitive.
In particular, most EA careers and most careers to do with AI are competitive as fuck.
Post 49: Things That Make Me Enjoy Giving Career Advice Said 17 Jun 2022
In any mentoring conversation, only a small fraction of what's discussed turns out to be highly valuable.
in general, I expect 90% of a mentoring chat to be kinda useless, and 10% to be particularly high value.
Post 49: Things That Make Me Enjoy Giving Career Advice Said 17 Jun 2022
In prioritizing a list of tasks, the closest calls matter least while getting the very top item right matters most
In hindsight, the hardest comparisons are in fact the least useful - if I do the fourth most important task before the third most important, that's no big deal, but doing the tenth before the first is a screw-up!
Post 48: Prioritise Tasks by Rating not Sorting Said 15 Jun 2022
Having any ordering to follow, even an arbitrary one, helps overcome indecision about where to start work
There's a surprising amount of value in having any order at all (even an arbitrary one like first-come-first-served!) to break past indecisiveness about where to start.
Post 48: Prioritise Tasks by Rating not Sorting Said 15 Jun 2022
Being smart and competent makes someone right more often, but not right every time.
Someone being smart and competent just means they're right more often, not that they're always right
Post 47: How I Formed My Own Views About AI Safety Said 27 Feb 2022
Mathematics gives practitioners corrective feedback on being wrong that moral philosophy lacks.
Mathematicians get feedback re whether there proofs work in a way that, as far as I can tell, moral philosophy doesn't
Post 47: How I Formed My Own Views About AI Safety Said 27 Feb 2022
Thinking harder about a decision has diminishing returns because reality is inherently unknowable.
Reality is not fully knowable. And thinking harder has costs.
Post 46: Reward Good Bets That Had Bad Outcomes Said 22 Feb 2022
Many highly intelligent people tend to be insecure and risk-averse.
many of the smartest people I know are super insecure and risk averse.
Post 46: Reward Good Bets That Had Bad Outcomes Said 22 Feb 2022
Pursuing a very low success-rate strategy can still be rational if a rare win is valuable enough.
It's obviously worth it to go on 99 unsuccessful dates if the hundredth results in marriage.
Post 46: Reward Good Bets That Had Bad Outcomes Said 22 Feb 2022
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
Feed
As its own page →Hiding
27 December 2025
19 August 2025
From one piece MATS Applications Open (Due Aug 29) 3 beliefs · neelnanda.io
-
Their words
My model is that research requires a mix of skills. The day-to-day coding and execution is crucial. But there's also a set of harder-to-learn conceptual skills, collectively called research taste. These skills take a long time to gain because they have poor feedback loops, but they take very little time to use.
-
Their words
I'm more agnostic about the best techniques, things like sparse autoencoders are a useful tool, but easy to waste effort using when a simpler method is sufficient or better - start by doing the obvious thing!
-
Their words
I'm now fairly pessimistic about ambitious interpretability (i.e. complete reverse-engineering), and I'm excited about model biology (studying qualitative high-level properties of models) and applied interpretability (rigorously doing useful things with interp).
26 May 2025
From one piece Post 51: Socratic Persuasion: Giving Opinionated Yet Truth-Seeking Advice 2 beliefs · neelnanda.io
-
Their words
It's a lot easier for someone to engage with an argument if they generated the key steps themselves by answering my questions - if imposed by me, it sparks contrarianism and defensiveness
-
korrents.com
The person receiving advice usually knows more about their own situation than the person giving it does.Their words
No matter how much I know about a domain, the other person will always know far more about their own situation, context, beliefs, skills, preferences, etc than I do.
Nothing matches.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.