A model looking like it is becoming smarter, attempting shenanigans less often, and more often doing what you want, but getting better at hiding its actions when it wants to do that, is exactly the scary combination.
Now, with Astra, we are no longer playing on super easy mode. The AI is going to think ‘will this obviously turn out super badly for me if I try it?’ and if the answer is yes then it won’t try to do the thing.
this is all about support for the ongoing Israeli campaign of ethnic cleansing and genocide, which continues with daily slaughter of civilians and demolitions of their homes and communities in Gaza, the West Bank, and now southern Lebanon.
there exists a wisdom in the involuntary, and that our attempts to shape reality, however intentional, are themselves a form of unintentional expression.
Agents will need persistent, unique identifiers that allow their actions to be traced back to a responsible actor. Doing this successfully will also require human users to possess a unique identifier.
I have said for many years now that the thing I'm worried about with the models is a Black Monday type scenario, where many algorithms work with each other and get us into weird basins of actions.
You have to persevere. And that's the other really important part of motivation. You know, we know a lot of people who start things and then we also know the non-competers, the people who are always so energetic but actually don't finish, right? So, it's almost as important to be able to finish and complete as initiate actions.
no matter what kind of industry you're looking at, there was this government angle where the government's policy actions could be ones that could actually unlock technologies.
So, you might be surprised to hear that most state-of-the-art foundation models for robotics have no memory or no context. They're just operating on the current sensor observations, the current camera readings, uh, and predicting actions based off of that.
If you think about our large scale models today, they probably see a thousand times as much data as a human does by the age of 18. Yet, the human by the age of 18 is better in a lot of things and, you know, on par uh with those frontier models that have seen way more data. So could you come up with much more data efficient systems that can learn continuously learn from their own actions?
So I think the the thing I'm most excited is actually like what we call like iterated loops or like slow loops where we basically have a cron job. We have the loop the the the structure of the loop is really easy. It's like run this llinter fix one thing commit and push and then we run that every night in our GitHub actions and we wake up every morning to one PR that makes the codebase a little bit better.
untrusted repos should be treated as hostile by default because they can steer the agent toward reading files, running commands, or sending data through approved tools.
You cannot just reverse that, right? It's just not working. So, I do think that we need to build out the whole guardrails for the reversibility of actions because that actually where things get really really scary.
behavior and belief is like two-way street. You know, what you believe will influence the actions that you take, but the actions that you take can also influence what you believe about yourself. And you know, every time you show up and do it in some small way, you prove to yourself a little bit, hey, maybe I am that kind of person. And so, my encouragement, my suggestion is to let the behavior lead the way.
are your small actions accumulating or are they evaporating? You know, are you doing small things each day that are oriented toward a larger outcome or are you doing small things each day that are just kind of like one-offs and don't really add up? In a lot of ways, I feel like the two time frames that matter most in life are 10 years and 1 hour
I’ve got a GitHub actions thing that runs a piece of software I wrote called shot-scraper that runs Playwright, that loads up a browser in GitHub actions to scrape that webpage and turn the results into JSON, which then get turned into an atom feed, which I subscribe to in NetNewsWire.
I’ve got a GitHub actions thing that runs a piece of software I wrote called shot-scraper that runs Playwright, that loads up a browser in GitHub actions to scrape that webpage and turn the results into JSON, which then get turned into an atom feed, which I subscribe to in NetNewsWire.
Like now, like basically learning is not for these systems is not just learning from raw actions. It's also learning from words eventually be learning from observing what people do from the kind of natural feedback that you receive when you're doing a job together with somebody else.
So there is an angle of, the US' actions, from the angle of the expert controls, have been so inflammatory at slowing down China's progress on the leading edge that they've turned around and have accelerated their progress elsewhere because they know that this is so important.
We ought to be especially careful in the cases where what we delegate to a device, app, agent, or system is an aspect of how we express care, cultivate skill, relate to one another, make moral judgments, or assume responsibility for our actions in the world
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.