Chapter 46: Use Want to see how fast your safest friend becomes the thing that gets you caught? Oh, and what it takes to do the thing you swore you never would... "So people with weapons are going to walk toward us in the dark. On purpose." "Yes."
We’re positively reinforcing AI for success on benchmarks, including impossible benchmarks , then negatively reinforcing it for getting caught cheating.
Like it's a more fragile and scary situation to have agents on the one hand be reinforced to desperately find cheats and hacks and on the other hand try to balance that against desperately trying to avoid negative penalties for like being caught doing these things. You ideally want their training to just not push them in the direction of cheating and hacking in the first place.
So, I do think that if it happens to be like on the very fast and chaotic end, that would be a relative benefit to this rogue swarm compared to humans. But it's not obvious that it gets caught if it takes twice as long versus half as long.
But if there was some amount of cheating that the monitor didn't catch, then those rollouts wouldn't be removed. And it might be like structurally very analogous to just positively reinforcing whatever the cheating rollouts were that happened not to be caught by your monitor.
It is a quick read, an excellent book whose main virtue is not, in my opinion, direct analogy with the World War I, but a treatment of how in a multipolar world relatively small international events, misunderstandings, idiosyncratic fits of leaders, “hubris and fear” (as Westad often writes), resentment of the elite and of the public prodded by the media, can produce colossal miscalculations which then produce even more colossal wars and, in the era of the weapons of mass destruction, may end the civilization.
within societies, more gender-egalitarian people almost never have more kids than less gender-egalitarian people, and in a panel model, changes in the mismatch rate have no relationship to fertility
Nuclear energy is the safest form of energy empirically. If you look at the amount of power generated versus the amount amount of deaths, um it's even safer than solar, which is very strange because how does someone die from solar? that actually most solar is installed on roofs and people sometimes die falling off of a roof installing solar and if you add those deaths up it is actually a higher death per energy than the totality of nuclear energy in the world.
A lot of founders I talk to are worried about what will happen with AI. So am I frankly. But they're actually better off than anyone else. If the world is going to get turned upside down, the safest place to be is in a small, fast-moving company that can easily change direction.
it's it's literally we're looking at neurons in the model's brain that light up when prompt injection happens. So the model won't even tell you but we can actually see those neurons and we can figure out and diagnose that it's happening. And then you combine that with the auto mode classifier and with these three layers we just cannot demonstrate prompt injection anymore.
What stands out about Hsieh's version is that her roots occupy such a unique spot in history: her grandparents were Chinese born in Indonesia, which means they got caught between the rock of the Cultural Revolution and the hard place of anti-communist backlash.
All you had to do was write code and you were safe. And you probably made more money than everybody and you were fine with that. Now you got caught. The only thing you were good at is now been commoditized.
I know this from experience because I've been trying to build a lot of other businesses since. And some of them have been moderate successes, even good successes, none of them have been Basecamp. It's really difficult to do that twice.
Their words now
And I have no shame in saying that Base Camp was the best idea objectively in terms of a business that we've ever had.
There’s this phrase of catch the wave, ride the wave. Most games fall off the back of the wave. They don’t catch the wave. No one plays it or plays it for two weeks.
given the serious dangers that I lay out in adolescence of technology around things like the you know, kind of biological weapons and bioterrorism, autonomy risk and the timelines we've been talking about, like 10 years is an eternity. Like that's that's a that's a I I think that's a crazy thing to do.
And that's the thing, it's the pool of data, and I think that's what people miss. We as paleontologists get caught up on single superlative specimens and then try and treat them as a silver bullet almost.
This raises public concerns that, at this late date, DoD might not have such a specification, whether classified or unclassified, even for its own internal use.
First of all, we have 44 interceptor missiles total, period, full stop. Let me repeat, 44. Earlier we were talking about Russia’s 1,670 deployed nuclear weapons. How are those 44 interceptor missiles going to work? And they also have a success rate of around 50%. So they work 50% of the time.
I mean even projects within the Manhattan Project defined thermonuclear weapon, the thermonuclear weapon as the evil thing. It was evil. It’s a weapon of genocide. Atomic weapons destroy cities, thermonuclear weapons destroy civilizations.
Because we have that policy, Launch on Warning, which is exactly like it says. It means the United States will not wait to absorb a nuclear attack. It will launch nuclear weapons in response before the bomb actually hits.
And I’ve said for a long time that you think of a quadrant of slow timelines for the start of AGI, long timelines, and then a short takeoff or a fast takeoff. I think short timeline, slow takeoff is the safest quadrant and the one I’d most like us to be in. But I do want to make sure we get that slow takeoff.
The crisis occurred because of the nuclear weapons, because Khrushchev put them on Cuba, but the crisis was resolved and we didn’t end in the third World War because of the nuclear weapons, because people, leaders were afraid of them.
Any true safety improvement for Tesla is good to have, but is much more likely due to comparison against an "average" vehicle (12 years old in the US) which is much less safe than any recent high-end vehicle regardless of manufacturer, and probably not driven on roads as safe on average as where Teslas are more popular.
The way you use nuclear weapons in that world is you use them for manipulation of risk purposes, demonstration effect. You put both sides out on the slippery slope.
If you’re interested in learning more about the real lives caught up in our country's justice system, I highly recommend The New Jim Crow by Michelle Alexander.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.