in the future, humans won’t find a bug or an issue with the code produced by a model, at least not in a reasonable time. Humans will only review the system and its composition, but it won’t be in PRs and it won’t be by looking through every line of the code.
Open source in its current form doesn’t make sense anymore. “Given enough eyeballs, all bugs are shallow” is still true but now we have artificial eyeballs
I guess our best defense right now is dependency cooldowns - giving new package releases a few days before upgrading to them, in the hope that supply chain attacks like this will be spotted by someone else.
@AntithesisHQ – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. https://t.co/AKYm4cbVCU
@AntithesisHQ – verify your system’s correctness without human review or traditional integration tests – and avoid bugs or outages. https://t.co/AKYm4cbVCU
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
Don't try to anticipate anything. You will literally go crazy. Because even the smartest brains in the business cannot anticipate what two model hops from here is going to look like. It's an absolute waste of time, and you will develop an AI psychosis trying to deduce what two years from now is gonna look like. Focus on right now, and right now is the most incredible time to be into computers.
Maybe I just read too much AI text now vs before, but I really feel like the model's ability to produce text I actually want to read and understand went downhill with newer releases.
Every step we take that makes the models less governable and controllable also makes them less useful, so we will not be able to continue using them, which reduces the value, and which means delaying the release is the only option
what what is the current thing that models are not very good at doing? Alpha is going to decay in some way with every model release or every series of model releases. So, it's going to change over time.
Facts and Fallacies of Software Engineering by Robert L. Glass. In essence, this is a book about an industry that refuses to learn. That was true 25 years ago when this book was published, and it's probably twice as true today. (Just think about all the AI adoption metrics being rolled out — back to productivity mistaken for lines of code produced, only more elaborate. And expensive). What I like about this book is that Glass doesn't present anything new. Quite the opposite, actually. Rather, it's about research lessons that we all should know, but tend to forget. Ever had to do an estimate, or plan according to a requirements spec? Or maybe you thought that enough eyeballs make all bugs shallow? Then this book is for you. A great work by a fantastic author.
Then the actual system might still have bugs, but we can iron out the issues in the abstraction such that we don't actually build them in the real system.
This study directly shows that working really hard, releasing a lot of cortisol in high-intensity interval training actually improves hippocample structure and function um over over time.
there's a difference between belief and hope. Hope is confidence without basis. Hope is just, you know, it's it's a prayer, but it's but it's it's not it's not founded in anything that your your lived experience.
I think part of what I liked about Rust is this feeling that as you write the code when it compiles it works. I mean this has to be in quotes right because obviously it's possible that there are bugs but this is something a lot of people say about Rust and there's a reason people say it even though it's not necessarily literally true.
Memory safety is this idea that no matter how stupid the code you write is, it's not going to have a certain class of bugs. And this is the, you know, the kind of bug that usually leads into security vulnerabilities.
I mean, that's the thing with memory safety, right? You make some sort of trivial mistake. It it's not some complex thing. It's just you make some trivial mistake and there's a bajillion places you could make it.
I think the same kind of principle applies with agents in that they can talk to the compiler. It will tell them what to fix. So I guess this could be a case. We we'll see. But Rust could be a pretty promising candidate for to use for agents because they can get more feedback and it's just hard it's harder to to ship certain type of bugs or maybe impossible to have certain type of bugs.
exceptions are an inherently poor way of handling errors because they make it easier to write bugs which won't be immediately obvious on casual code inspection
And one of the things I noticed is for 25 years we've kind of had the answers. Somebody comes to us and says we have too many bugs or like, all right, well, here's how you write tests. Oh, I can't write tests. Well, here's how you design so you can write tests. It's just kind of press play uh on the recorder. And the thing that's changed is at this moment nobody knows the answers to anything.
You know, obviously, we patch our games and that’s where we fix a lot of bugs, but if you really wanna run a game like Overwatch or World of Warcraft successfully, you need master level engineers who have architected the client and server in such a way that you can hotfix the game on a dime.
I switched to Firefox when Safari 15 was released. John Gruber summarized the issues I had with it perfectly. Since Apple reverted their design changes in recent releases, I'm back at using Safari as my primary
I switched to Firefox when Safari 15 was released. John Gruber summarized the issues I had with it perfectly. Since Apple reverted their design changes in recent releases, I'm back at using Safari as my primary
Well, what we realized really early is that quantity of employees doesn't translate the quality of the product they produce. In many cases, it's the opposite. If you have too many people, they have to coordinate their efforts, constantly communicate, and 90% of their time will be spent on coordinating the small pieces of work they're responsible for between each other.
Bigger changes always get tests. Automated ones usually aren’t great, but the model almost always finds issues when you ask it to write tests IN THE SAME CONTEXT. Context is precious, don’t waste it.
The US will lead for some time on breakthroughs on disruptive technologies, the zero to one technologies that ultimately change the world. But innovation is a process. It goes from invention to production and commercialization and diffusion, diffusing technology throughout all parts of the economy. And on those two stages, I think that China has a unique advantage, even if it still can’t do the zero to one breakthroughs, because in the end, how much this technology is adopted by the countries and by the various parts of the economy is fundamentally crucial to how much productivity will be unleashed.
For me personally, I kept introducing bugs, and I couldn't figure out why. And what I realized is that I developed... I wasn't copiloting well, I was autopiloting much better.
For now, I am not convinced that issuing press releases about your compounds that talk about their discovery through AI techniques is sufficient to expect greater things from them.
Managers should support 6-8 engineers. This gives them enough time for active coaching, coordinating and furthering their team’s mission by writing strategies, leading change, and so on.
This book had a big impact on me as young software developer. It was the first time I saw in print the principle "bugs that go away by themselves come back by themselves". It's something that I've often repeated to myself, and my team, if nothing else but to prepare us mentally for when that hard bug does come back. The chapter advising developers to "proactively step through your code during development, don't wait for the bug report" improved my productivity enormously, and I followed the practice for many years.
I use all the Web browsers a lot since I'm constantly chasing browser bugs, and don't currently really like any of them though I lean towards Chrome (but currently am annoyed at it because it keeps crashing on me).
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.