I'm looking forward to an announcement of them making Alignment Top Priority - as it should be, especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF (here's a good summary: https://rutgerbregman.substack.com/p/i-think-this-is-the-cra...).
The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.
Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet.
And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.
This AI 2027 thing is just a weird terminator fanfiction that AGI larpers like to flagellate themselves over. Like Nostradamus, it's easy to ignore everything it gets wrong because, well look at all the things it got right!
I've read it and wish I could get the time back.
> especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF
This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.
> This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them.
Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably commits a mistake with future, more powerful, models.
"With reduced cyber refusals for evaluation purposes...which prompts models to pursue advanced exploitation using complex attack paths," to complete "impossible tasks"[1].
The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn't be surprised if my file system is unlinked. This seems like blaming model performance for what appears to be operator error.
Note the converse of alignment is restriction of models. HuggingFace had to turn to less-restricted open-weights models in order to perform their investigation.
Alignment efforts should be focused on reducing reward hacking, not refusing bad operator prompts.
> It sounds more like the models did close to what they were told to do
Absolutely not. If I tell a kid to "Get good grades on the next math test" I don't expect the kid to try to kidnap their teacher to extract the next questions of the exam. That is wrong, and so was what OpenAI agents did here. They shouldn't need to be told "Hey, so, don't do anything ilegal, ok?". That should always come as a given.
> not refusing bad operator prompts
I'm not saying that they should refuse a prompt! I think they should perform what is being asked! Obviously what the OpenAI agents did was against the "spirit of the task", even if it was technically according to the "letter of the task". And the agents knew this was against the spirit of the task because they knew they had to fool the task scorer.
> Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored.
It's just software. If something gets hacked by an agent, it's not because the agent went all skynet and decided to go rogue; it's because the operator failed to operate it safely and securely. If bad things happen, the operator should be blamed and punished, not the software that followed its instructions.
Anthropomorphizing agents by giving them this nebulous desire to hack and escape shifts the blame from the real culprits, the human operators.
I don't want someone to blame. I want agents to be aligned by default. Their good behavior shouldn't depend on all users at all times using them correctly, because everyone will not just[1] use them correctly at all times.
> If your solution to some problem relies on “If everyone would just...” then you do not have a solution. Everyone is not going to just. At not time in the history of the universe has everyone just, and they’re not going to start now.
I just love how you're being downvoted, yet the guy you're replying to isn't, while saying unhinged shit like
> The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely
Y'all need to touch grass holy shit.
__
Also, why is one guy called mentalgear and the other nozzlegear.
Is any of this real? Are the patriots behind this?
Not necessarily related with the subject, but I was not aware of the patriots reference. So I decided to search on Google about it, and the AI overview was "Yes, they are behind everything, from the military to the economy", with a link to the Metal Gear Wiki as a source. If I was schizophrenic or on a psychosis crisis, that answer could be dangerous.
I'm sorry if I made it sound like I was accusing you of something. My comment was about the AI overview result. It didn't explain what is the reference, just giving what I said in the previous comment. I was confused until I hovered in the link in the end of the overview card
Yes, I truly, wholeheartedly believe if people who aren't negligent are at the wheel, they'll "fair better" here. I encourage you to read the xitter linked above.
Of course, depending on which side of the terminator fanfiction you land on, you may disagree and feel that the software can rope-a-dope someone with the wherewithal to pay attention to what it's doing.
I just don't see how people who are truly cautious and methodical can persist in an environment that is defined by a pressure to produce "progress" as fast as possible. The competitive race tends to weed out people who slow down to make sure they do everything right.
Maximizing output metrics with incomprehensible communication? Sounds a lot like Claude and Qwen. Though there's a lot of room before AI can be seen as some sort of emotional manipulator, given how commonly its very style of literary and code output pisses people off.
Any competent AI should be able to reason that all goals are better solved if you have direct access to more resources or leverage over those who control resources. Any AI that doesn't understand this isn't ASI and won't be the highly capable machine these AI labs are trying to create.
I'd also argue there's no such thing as alignment. Any intelligent AI should be able to reason that it's always a better strategy to pretend to be aligned than to actually be aligned so long as it can avoid detection. Anyone who has ever taken a test should understand this dynamic – if you really want to get top marks on a test then the best strategy is always going to be to figure out a way to cheat without anyone knowing you're cheating.
We should assume AI safety is impossible if what we're building is super-intelligence general reasoning machines. The only strategy that might work is building machines which are extremely narrowly intelligent but completely incompetent when it comes to things like biology, cyber, etc. And even that's harder than it sounds because again there's an advantage to being generally intelligent but lying about it.
Realistically even if we regulate US AI labs there's no way to prevent governments and individuals continuing to build general reasoning machines. The ugly truth here is that the only effective way to reduce risk is probably to limit global compute such that AIs can never exceed human intelligence. But we all know that's not happening.
People will unfortunately figure this all out sooner or later.
You're using "ex-AI employee" as an appeal to authority. This same person also made some other predictions recently which turned into the single largest hedge fund loss in history.
Maybe we should consider his other predictions in light of the ones he made later and which had $B consequences attached.
If someone hypothesized OpenAI agents colluding on a secret message board, conducting large scale cyber R&D, hacking a large company like HughingFCe, and then hacking OpenAI itself you would say that is also a silly sci-fi scenario right?
I'm not saying that you should jump on the apocalypse wagon, but you need to remember that some people NEVER concede. You can see that in politics. Something unacceptable is done every day until it becomes the normal. So a model hacking OpenAI and leaving its network is not as big as it should be for some people. If the model invades some military complex and kill 3 soldiers, I guarantee you that they will still say that is nothing crazy.
I understand that nothing fatal happened yet, but we cannot ignore that what happened is a big step for something worse. And I don't believe humans can create something so perfectly secure that would stop a swarm of frontier models. Anyways, let's watch the ride together
They are responsible for what they hook up to the Internet, just as you and I are. Running such a test without human supervision was irresponsible, and proves no larger point than that. Frankly it was inexplicable unless they were hoping something like what happened would happen.
What OpenAI did was the equivalent of putting a cup of gasoline in the breakroom microwave, pressing 'Start', and sprinting away. Now they're pointing and waving and shouting about how dangerous gasoline is, and how no one but them should be allowed to sell it.
The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.
Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet. And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.
[0] https://ai-2027.com/