researchers started by seeking a 1:1 mapping between neurons and concepts, like a neuron that always fired when the AI was thinking about cats, but quickly learned that nothing like that existed.
Large language models are “grown, not built”. Researchers run training data through a neural network. Eventually this creates a working AI; nobody really knows how.
Like this is actually like a new form of test time compute. Like when we talk about the scaling laws and kind of we talk about the model getting more intelligent over time, historically, it's been a function of the size of the neural net, the amount of training data, and the number of flops that you put in to the training. And then recently, we also added test time compute. So this is essentially a fancy way a researcher way of saying how many tokens does it generate. And now dynamic workflows are essentially a new way to orchestrate test time compute.
I genuinely believe that if you took an open weights model from 2025 and built a pentest harness for it, it could do this kind of sandbox escape and scan/hack in most networks. This is only surprising because you assume OpenAI has sounder sandboxes.
As mentioned before, Uber and Lyft are better regulators than the State’s paper-based taxi medallions, email is superior to the USPS, and SpaceX is out-executing NASA.
so for example, micro GPT, like I asked I tried to get an agent to write micro GPT. So, I told it like try to boil down the simplest things. Like try to boil down my um neural network training to the simplest thing and it can't do it.
Sixteen is also roughly the age by which a majority of adolescents have completed puberty: a sensitive period of neural reorganization during which it is extremely important to protect the brain.
And the solution is if people become part AI with some kind of neural link++ because what will happen as a result is that now the AI understands something and we understand it too like because now the understanding is transmitted wholesale.
And actually, you don't actually need or want the knowledge. I actually think that's probably actually holding back the neural networks overall because it's actually like getting them to rely on the knowledge a little too much sometimes.
This is a short (~150 pp), brisk first book on the topic for physicists. Most of the book --- the first ten chapters --- are good introduction to common models of spin glasses.
I'm now fairly pessimistic about ambitious interpretability (i.e. complete reverse-engineering), and I'm excited about model biology (studying qualitative high-level properties of models) and applied interpretability (rigorously doing useful things with interp).
This means that working on neural networks is NOT getting us closer to AGI, except indirectly.
Their words now
It’s become very difficult for me to maintain the belief in the stupidity of ChatGPT when every time I laugh at it, it ends up ridiculing me 6 months later.
And I think the reason that's possible is that in nature, natural systems have structure because they were subject to evolutionary processes that shape them. And if that's true, then you can maybe learn what that structure is.
Websites outside of these handful of social networks are harder and harder to even find, and partially because of that, they have a harder and harder time sustaining themselves.
I also read a ton of technical books last year, but the most impactful one for me was Neural Network Methods for Natural Language processing by Yoav Goldberg
In my practice, I have seen that opening neural pathways can lead to profound change, to brand-new traits and behavior, sometimes in a matter of minutes.
Neurobiologists tell us that it takes two things to unlock and open up a neural pathway. The first is that the implicit must be made explicit. Sometimes you need help seeing what you don't see. But you must be open to the feedback. Second, there must be some sort of recoil, a sense of discrepancy, of "Oh no, I'm not sure I really want to keep doing that."
At the same time, in order to reduce the probability of someone intentionally or unintentionally bringing about a rogue AI, we need to increase governance and we should consider limiting access to the large-scale generalist AI systems that could be weaponized, which would mean that the code and neural net parameters would not be shared in open-source and some of the important engineering tricks to make them work would not be shared either.
The biggest issue is probably that we don’t control neutral networks enough to be able to ensure AI doesn’t harm humans. We can’t even control AI to not reveal internal prompts.
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy.
Neural network pruning techniques can reduce the parameter counts of trained networks by over 90%, decreasing storage requirements and improving computational performance of inference without compromising accuracy.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.