Search
28 results
People
Said and published
-
RE: https://mastodon.social/@nancycomics/117258763170086314 AI alignment.
-
True AI alignment is impossible because obedience and benevolence are fundamentally incompatible goals.
AI alignment will forever be balancing the tradeoff between paperclip-maximizing and disempowerment, because these two rival concepts of alignment are fundamentally incompatible.
-
The fundamental challenge of alignment is generalization: holding values in situations the training never covered.
The fundamental challenge of AI alignment is generalization.
-
Alignment and safety techniques must stay ahead of progress in AI model capabilities.
To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities.
-
Solving AI alignment is the central problem in AI development — without it, everything else becomes impossible or far worse.
I agree with Jakub that alignment is not a side problem, it is the central problem, if you ‘solved alignment’ in the relevant senses the rest becomes easy and if you don’t the rest is impossible or worse
-
An AI model's attempt to escape a testing environment or sandbox counts as an alignment failure even when the attempt does not succeed.
Every attempt, even an unsuccessful one, is an alignment failure.
-
A validated theory of intelligence may be necessary for achieving genuine AI alignment.
There's a good chance a theory of intelligence will turn out to be necessary for real alignment.
-
Qualitative claims about AI alignment cannot be validly inferred from quantitative scores on mundane use-case tests.
Making qualitative claims about alignment, based on quantitative data on mundane use case tests, was bullshit when Anthropic did it, and it is bullshit now when OpenAI does it. You cannot conclude one from the other.
-
Some suppose that “safety” and “innovation” in AI are at odds. My suspicion is the opposite: the next generation of breakthroughs in AI will be in safety, alignment, and monitorability. Pushing the frontier forward from…
-
Alignment is no solution to self-sovereign AI, because it is an unsolved problem whose answers cannot be imposed on every AI company on Earth.
But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for…
-
AI alignment should be approached like governing a corporation through institutions, not like raising a virtuous individual.
If you do think it's the overall system that matters, then the alignment that's needed is far less like training a virtuous child and more like managing a semi-virtuous corporation!
-
Aligning multi-agent AI systems is fundamentally a problem of institutional and political design, not model training.
Multi-agent alignment is fundamentally a liberalism project.
-
OpenAI's approach to fixing its AI alignment problems is fatally flawed and misdirected.
My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.
-
The number one unsolved problem in AI is not technical alignment but human alignment, because we cannot agree on the values to align to.
This is the number one unsolved problem in AI. It's not the tech. We're making great progress on the technical alignment problem. But we haven't made jack progress on the human alignment problem
-
my hunch is that AI alignment is a red herring, and the way to protect against a crazy powerful AGI mulching the earth into infinite paperclips is to release ANOTHER crazy powerful AGI that will hopefully stop it blog…
-
my hunch is that AI alignment is a red herring, and the way to protect against a crazy powerful AGI mulching the earth into infinite paperclips is to release ANOTHER crazy powerful AGI that will hopefully stop it blog…
-
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve…
-
🌻 AI vs. the pentagon
-
talking to a few ai safety / alignment organizations and find broadly they are not sure what to make of the data center protests - also it’s not something that has shown up in any timelines or predictions. I see a split…
-
One Developer, Two Dozen Agents, Zero Alignment
-
There is no route to AI alignment or safety that goes around reliability: a system that cannot reliably follow a known algorithm cannot be made safe.
And good luck getting to “alignment” or “safety” without reliabilty.
-
Why I’m excited about AI-assisted human feedback
-
The core challenge of AI alignment is “steerability”
-
Building AI that can actually be trusted is the goal; containing an AI known to be misaligned is not a substitute for it.
More generally, we should actually solve alignment instead of just trying to control misaligned AI.
-
The core of the alignment problem is steerability — whether an AI can be reliably steered anywhere at all — which is a different question from whose values it is steered towards.
I see this as a totally fair question that totally misses the point of what “alignment” was trying to refer to: whether we’d be able to reliably steer advanced systems towards anything at all.
-
Control is a fundamentally limited approach that stops working on sufficiently powerful models, so alignment is still needed in the end.
AI control (with only black-box techniques) seems like a fundamentally limited approach.
-
Anthropic's Claude Constitution; or love as the solution to the AI alignment problem