Related posts
Lilian Weng Blog
The the words it uses that this site has seen least often elsewhere. Posts are matched on those words alone — nothing here is a summary of this one.
Top people
Showing Profile →
Hiding
Hiding
20 September
19 September
18 September
17 September
-
korrents.com
Jev is basically a general purpose classifier.Their words
jev is basically a general purpose classifier. and i think that is amazing.
5 September
1 September
-
Their words
So, we did a classifier sweep specifically looking for agents thinking about or making the decision to alert humans. And across, 1200 transcripts, each of which are extremely long, we only found like a halfozen instances of it ever occurring to any agent to potentially notify humans. Um, and all of them just decide not to do it.
15 August
-
Their words
the score is the classifier's estimated probability for the AI-generated class based on its training distribution. However, we shouldn't interpreted it as a general probability that the text was written by AI.
Building an AI Text Detector From Scratchmagazine.sebastianraschka.com
-
Their words
if you look at the nuclear industry outside of Valor, it's mostly a modeling and simulation uh industry. Like when you think of a nuclear company, right, there's like nuclear companies out there that everyone knows the name of. And you kind of look under the cover. It's actually a modeling and simulation company, right? They produce um very very precise what we call paper reactors which have really good predictions of how a theoretical thing might behave.
3 August
-
Their words
Turns out, the task of driving is not that dissimilar from the task of modeling language uh because of the social aspects of driving. You're kind of having a conversation with other dynamic actors in the world, but you're doing that in the space kind of body language of your agent, your car, as opposed to just the language of words.
27 July
-
Their words
it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.
16 July
1 July
2 June
6 February
23 January
-
korrents.com
A personalized classifier needs several hundred labeled examples before its predictions become reliable.Their words
Predictions were poor until I got to over 250 samples in the training set.
Email triage with an embedding-based classifieradamwiggins.com
14 January
21 January 2025
19 September 2024
-
korrents.com
Biology keeps turning up whole classes of things nobody knew existed, and that is what defeats attempts to model it.Their words
The "unknown unknown" problem makes hash (in retrospect) of many attempts at biological theorizing and modeling, and I think that this problem will still accrue to ML models until we know an awful lot more.
18 January 2023
4 March 2022
-
korrents.com
The next-token language modeling objective is misaligned with following user instructions helpfully and safely.Their words
This is because the language modeling objective used for many recent large LMs-predicting the next token on a webpage from the internet-is different from the objective "follow the user's instructions helpfully and safely" (Radford et al.,, 2019; Brown et al.,, 2020; Fedus et al.,, 2021; Rae et al.,, 2021; Thoppilan et al.,, 2022). Thus, we say that the language modeling objective is misaligned.
8 February 2022
-
Their words
But I do want to stress that each of these five approaches are all just trade-offs, and none of them are universally wrong.
Nothing matches.
What is a korrent?
A korrent is a belief a person has stated in their own words: one sentence stating the claim, backed by a quote and a source, kept at korrents.com.
Under a name here, the quoted block is what they actually said. The korrent beneath it is the claim those words support, in korrents' wording — tap it to see the record, its source, and who else holds it.
Nobody here wrote their own korrents. They are compiled from public statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they do, this site shows a machine translation beneath the post, in this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the serif above is exactly what the person published, and it is what to quote them on. A translation can be wrong in ways that matter, especially about tone.
Only the post's own words are translated. A quoted post, a linked article and a belief on korrents.com are left in their original language.