From one piece Richard Sutton – Father of RL thinks LLMs are a dead end 5 beliefs, in the piece's order there
-
Their words
A lot of it has to do with just how you feel about change. Um, and if you think the current situation is really really good, then you're uh more likely to be suspicious of change and averse to change than if you think um it's imperfect. And I think it's imperfect. In fact, I think it's pretty bad.
-
Their words
It's our choice whether we should say oh they are our offspring and we should be proud of them and we should celebrate their achievements or we should we could say oh no they're not us and we should be horrified.
-
Their words
We don't we don't really know what information they had prior. We are we have to guess because they've been fed so much. This is one reason why they're not a good way to do science. Uh it's just so uncontrolled, so unknown.
+ 2 more
-
Their words
The scalable method is you learn from experience. Um you uh you you try things, you see what you see what works. No one no one has to tell you. First of all, you have a goal. So without a goal, uh there's no sense of right or wrong or better or worse. So large language models are trying to get by without having a goal or a sense of better or worse. That's just, you know, it's exactly starting in the wrong place.
-
Their words
reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.