In this case, however, the agents did not copy their weights, attempt to procure replacement compute, or take other steps that would be rational to take if their objective was to survive shutdown. So while the agents in the OpenAI-Hugging Face Incident were rogue, they were not truly sovereign.
when AIs are extremely extremely capable my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI. We do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand. Then we like can like go look in training and be like, "Oh, the these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there." But in a regime where the AIs are extremely situationally aware, very very very very capable and um you know uh we don't necessarily understand what they're doing, this feedback loop breaks down.
reliability and performance lives on this exponential ladder of nines. So, getting to that first 90% or 99%, that's the easy part. But then every next nine that you want to add, that takes about 10 times more effort.
Know where the money is coming from. Know where you plug into that. And like if you're just asking those questions, you're actually already like steps ahead of all of the other like fledgling perspective mathematicians.
In terms of science, I think it really makes sense and we're deeply committed to open source. Um, there are obviously interesting considerations on this that are important too because there's a lot of considerations around biosafety and things like that that we're going to need to balance and think through how to how to handle.
You probably only have one good idea in your whole product. Get the user to experience that as fast as possible. Well, it sounds so easy, but if you go try every single product out there, you will see a million accidental steps they introduced from the user hearing about your product to like seeing the value in it.
I think that, you know, the model that we've put together collectively about the relationships between archaic and modern humans is sort of accreted over time. There was this, you know, idea that modern humans are distinct and that Neanderthals and Denisovans are like sisters of each other. And then over time we developed and detected additional mixture events like this modern human into Neanderthal and then this other ones I didn't even talk about like super divergent lineage going into Denisovans and like all this other stuff.
I mean, I I I I do believe that that hybrid um human plus AIs will will dominate mathematics for a lot longer. It it's It will depend It will require some additional breakthroughs uh beyond what we already have.
My sense is that this hybrid is fairly common in practice; solvers aren't magical and if you can deduce additional structure using domain-specific analysis, it will often give the solver an important boost.
I think basically what takes the long amount of time and the way to think about it is that it's a march of nines and every single nine is a constant amount of work.
when you learn to play chess you have the grand the long-term goal is winning the game and yet you you can't you um you want to be able to learn from shorter term things like you know taking the your opponent's pieces um and so you do that by having a value function which predicts the long-term outcome
I I like the term context engineering because I think the fundamental skill of using AI well is to be able to state a problem with enough context in such a way that without any additional piece of information, the task is plausibly solvable.
I hope we'll end up with something more collaborative if needed, more like a CERN project where it's research-focused and the best minds in the world come together to carefully complete the final steps and make sure it's responsibly done before deploying it to the world.
It's a lot easier for someone to engage with an argument if they generated the key steps themselves by answering my questions - if imposed by me, it sparks contrarianism and defensiveness
But what we actually found was that none of those are actually hard. The whole idea of hard steps, that there are hard steps, is actually suspect. What's amazing about this model is it shows how important it is to actually work with people who are in the field.
An eye-opening view on considerations going into building a widely used public API or reusable library. While the book focuses on the .NET framework, many of the conventions apply to maintainable and reusable components, in general. This book had an outsized impact on me as I read it when I was a mid-level .NET developer.
And as responsible scientists, we’re trying to disprove our theories. We are not supposed to be trying to prove our theories. That’s one more foot out of the science box that archeology often steps.
Another challenge for Class II biotechs is that even if you have a nifty computational drug discovery platform, you are still often bottlenecked by the rest of the drug development process. It takes 1,000 (or more!) small steps to advance a drug through discovery, development, and past the regulators.
The principle in Perplexity is you’re not supposed to say anything that you don’t retrieve, which is even more powerful than RAG because RAG just says, “Okay, use this additional context and write an answer.” But we say, “Don’t use anything more than that too.” That way we ensure a factual grounding.
And what Lee’s been able to show in the lab with his team is that for organic chemistry, it’s about 15 steps. And then you only see molecules that the only molecules that we observe that are past that threshold are ones that are in life.
I liked Mem’s editing experience within notes because it resembled word processing, without extra steps getting in the way. But I found the app otherwise surprisingly lacking in basic functionality, despite its polished appearance.
Society's response, despite promising first steps, is incommensurate with the possibility of rapid, transformative progress that is expected by many experts. AI safety research is lagging. Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems.
The third UI paradigm, represented by current generative AI, is intent-based outcome specification. The user tells the computer the desired result but does not specify how this outcome should be accomplished, such as the steps to be executed. Compared to traditional command-based interaction design, this completely reverses the locus of control.
If my class is so boring that I have to devise an additional carrot-and-stick system to get students to pay attention, the problem is my class is boring.
Taken together, there is an argument to be made that AVs should be safer than human drivers by about a factor of 10 (being a nice round order of magnitude number) to leave engineering margin for the above considerations.
Code has mass. Every additional line of code you don't need is ballast. It weighs your codebase down, making it harder to steer and change direction if you need to.
This is the book that started the Lean Startup movement. While it is not an easy read, it is packed full of good information and helpful charts and guides.
Selling to big companies is a different process than selling to an individual. If you're new to sales, this book is a great way to understand the added steps and what to do. I loved how actionable the book was. You can literally open it up and know the tactic to use, the questions to ask, or pitfalls to avoid for each step in the process.
This is Steve Blank's textbook to starting a company. It's the Four Steps to the Epiphany, expanded and much more readable. If you want to deep dive into the Lean methodology, this is the book to read.
A korrent is a belief a person has stated in their own words: one
sentence stating the claim, backed by a quote and a source, kept at
korrents.com.
Under a name here, the quoted block is what they actually said.
The korrent beneath it is the claim those words support, in
korrents' wording — tap it to see the record, its source, and who
else holds it.
Nobody here wrote their own korrents. They are compiled from public
statements, and a person can change their mind, which is recorded too.
About the English under a post
Some people here publish in a language other than English. Where they
do, this site shows a machine translation beneath the post, in
this typeface — the site's own, not theirs.
The post itself is never changed, moved or hidden: what is set in the
serif above is exactly what the person published, and it is what to
quote them on. A translation can be wrong in ways that matter,
especially about tone.
Only the post's own words are translated. A quoted post, a linked
article and a belief on korrents.com
are left in their original language.