Lilian Weng
Everything, newest first — across every channel. Their profile →
Hiding
14 July
4 July
From one piece Harness Engineering for Self-Improvement 3 beliefs, in the piece's order there
-
Their words
harness improvement enables better deployment of the model but intelligence is still the core.
Harness Engineering for Self-Improvementlilianweng.github.io
-
korrents.com
Recursive self-improvement only works if the base model is already capable enough to improve its own mechanism.Their words
Recursive structure alone is not enough. The base model must be capable enough to improve the mechanism.
Harness Engineering for Self-Improvementlilianweng.github.io
-
korrents.com
The deployment layer surrounding a model is as important as the model's raw intelligence.Their words
the layer between the raw model and the real-world context seems to be as important as the model's raw intelligence
Harness Engineering for Self-Improvementlilianweng.github.io
24 June
-
korrents.com
As models keep growing, AI is running out of enough high-quality unique training tokens to keep up.Their words
As the model size grows significantly, we are running out of enough high-quality unique tokens.
1 May 2025
From one piece Why We Think 3 beliefs, in the piece's order there
-
Their words
model CoTs could be biased due to lack of explicit training objectives aimed at encouraging faithful reasoning
-
korrents.com
Longer test-time thinking improves an AI model's robustness to adversarial or unusual inputs.Their words
thinking for longer should be especially useful when the model is presented with an unusual input, such as an adversarial example or jailbreak attempt
-
korrents.com
Large language models cannot reliably self-correct their own mistakes without external feedback.Their words
this self-correction capability turns out to not exist intrinsically among LLMs and does not easily work out of the box, due to various failure modes
28 November 2024
From one piece Reward Hacking in Reinforcement Learning 2 beliefs, in the piece's order there
-
korrents.com
More capable AI agents are more likely to find and exploit flaws in their reward functions.Their words
A more intelligent agent is more capable of finding "holes" in the design of reward function and exploiting the task specification-in other words, achieving higher proxy rewards but lower true rewards.
Reward Hacking in Reinforcement Learninglilianweng.github.io
-
Their words
Reward hacking exists because RL environments are often imperfect, and it is fundamentally challenging to accurately specify a reward function.
Reward Hacking in Reinforcement Learninglilianweng.github.io
7 July 2024
From one piece Extrinsic Hallucinations in LLMs 2 beliefs, in the piece's order there
-
korrents.com
Using supervised fine-tuning to teach a language model new knowledge risks increasing its hallucination rate.Their words
These empirical results from Gekhman et al. (2024) point out the risk of using supervised fine-tuning for updating LLMs' knowledge.
-
korrents.com
A language model avoids hallucinating only by being factual and admitting when it does not know the answer.Their words
To avoid hallucination, LLMs need to be (1) factual and (2) acknowledge not knowing the answer when applicable.
12 April 2024
From one piece Diffusion Models for Video Generation 2 beliefs, in the piece's order there
-
korrents.com
High-quality video data, especially paired with text, is harder to collect at scale than image or text data.Their words
In comparison to text or images, it is more difficult to collect large amounts of high-quality, high-dimensional video data, let along text-video pairs.
-
Their words
It has extra requirements on temporal consistency across frames in time, which naturally demands more world knowledge to be encoded into the model.
5 February 2024
-
korrents.com
High-quality data collection depends more on careful human execution than on machine learning techniques alone.Their words
Lots of ML techniques in the post can help with data quality, but fundamentally human data collection involves attention to details and careful execution.
25 October 2023
From one piece Adversarial Attacks on LLMs 2 beliefs, in the piece's order there
-
korrents.com
Universal adversarial trigger attacks are easy to detect because the learned trigger tokens tend to be nonsensical.Their words
One drawback with UAT (Universal Adversarial Trigger) attacks is that it is easy to detect them because the learned triggers are often nonsensical.
-
Their words
High perplexity makes an attack more vulnerable to be detected and mitigated.
23 June 2023
15 March 2023
27 January 2023
10 January 2023
8 September 2022
9 June 2022
15 April 2022
20 February 2022
5 December 2021
24 September 2021
11 July 2021
31 May 2021
22 March 2019
14 February 2014
3 August 2012
2 August 2011
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.