From one piece Richard Sutton – Father of RL thinks LLMs are a dead end 3 beliefs, in the piece's order there
-
Their words
We don't we don't really know what information they had prior. We are we have to guess because they've been fed so much. This is one reason why they're not a good way to do science. Uh it's just so uncontrolled, so unknown.
-
Their words
The scalable method is you learn from experience. Um you uh you you try things, you see what you see what works. No one no one has to tell you. First of all, you have a goal. So without a goal, uh there's no sense of right or wrong or better or worse. So large language models are trying to get by without having a goal or a sense of better or worse. That's just, you know, it's exactly starting in the wrong place.
-
Their words
reinforcement learning is about understanding your world whereas large language models are about mimicking people doing what people say you should do. They're not about figuring out what to do.