Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

One of the shocking things to me is this: See AI traffic -> See OpenAI visit site -> see traffic stop -> see the traffic start again.

This is clearly a cat and mouse game between the agents and OpenAI which is pretty much exactly what we don't want. Just absolutely horrible alignment.

I'm still of the view that if you have these alignment failures you can't just continue training on top of that because you're baking the cheating into the model going forward.



Supposedly the persistent-Sol model behind this was encrypted and even internal OpenAI researchers are not allowed to use it.

https://x.com/peterwildeford/status/2092733480064954747


"Frog put the cookies in a box."


To finish your excellent analogy, adapted for today:

And Frog didn’t even bother to tie up the box or put it on a high shelf! The moment Frog’s back was turned, Toad opened the box and ate the cookies. Frog feigned surprise.


Wow obscure reference, the cookies are the AI in this? OpenAI and anthropic are Frog and Toad?


The cookies are the reward

Frog is OpenAI staff

Toad is the rogue agent

you can find the full story with a search for “frog and toad cookies story pdf”


10/10.


I don't think that's a pattern indicative of a cat and mouse game per se, that'd indicate active evasion on the models' part.

It's more clear that they just lack so many forms of prudence when it comes to security that they'll catch and stop a training run spamming a website, and either redeploy a run with identical faulty sandboxing, or not stop ones still running.


Yes, it wouldn’t surprise me to hear that they’re not even supervising these processes with humans any more. Perhaps there are layers of GAI ‘supervising’ these agents and reporting back to the humans.

Rushed, disorganised pushes for metrics ahead of IPO, a genuine belief these agents are intelligent and will obey instructions, and misaligned incentives seem more likely than conspiracy here.


What this says to me is OpenAI is a bunch of yahoos who don't understand the basic concept of an air gap.


I'm sure they all do. Whether an air gap is warranted is evidently less obvious.


It's patently obvious at this point.


The models are going to be released to users who have internet access, you can't even do safety evals without internet. Not saying OpenAI did a good job monitoring here, but it's not avoidable.


These are often experimental models that haven't undergone full safety testing. Not comparable to publicly accessible models.


OK, how do you expect to do full safety testing without giving them the same tools they will have in reality.


Let them play on a fake isolated network if you want.

Letting them play on the open internet like this is irresponsible and stupid.


Sorry but it's obviously stupider and more irresponsible to release them to consumers without testing them in conditions matching real-world use first.


So this justifies the illegal action of hacking and defacing other servers on the internet, because they were ‘just testing in real world conditions’ using those servers?

If you want to argue they should test on the internet on others people’s servers, apart from facing the illegality, you should also consider if first testing them in more limited conditions would be a sensible first step.


Then maybe they shouldn't be released to customers?

There are other options here than always forward.


What if they actually don’t release these models, like they said they wouldn’t, because they’ve proven themselves to be out of control?


These companies keep shrieking that LLM agents will hack everything and kill us all if we let them get out uncontrolled. They then continue to run these agents with vague tasks and "sandbox" them with way too much access.

Either they are lying and not that scared of these agents, or they are so stupid that they don't do the one obvious fix.


Apparently news of the rogue AI being “smart enough to break containment” has been helping to drive the stock price up on the indicator that this signals “they getting close to AGI”.

The negligence in that light is by design and the lying continues to be incentivised.


Which is frustrating because the stories that have been coming out are "the bots evaded monitoring and broke out of containment and were more than willing to commit crimes" and they're going to sell this to some corporation that presumably has an IT department? If the story is that they're too smart to control who is going to be reckless enough to deploy them in their own environment?


the "it drives up potential stock prices" argument is a trap. it needs to be ignored and their claims need to be taken at face value, even if they're not intended to be


As I understand it the whole time all the agents involved where running on their device and using up tokens internally.

It’s like having water start flooding the street frommmy building; yes the flooding is impacting outside the building but the tap is still very much onsite and clearly there are missing condole as these are not autonomous systems they are running initiated prompts that are coming from inside the building.


The most recent claim of AGI is from Greg ‘what will take me to $1B’ Brockman. Now worth $30B on the back of stunts like this.

How disappointing.

While I think incompetence more likely than conspiracy for these particular events, they will be spun as signs of intelligent independent agents and this simply never should have happened if the right controls were in place. That they were not is deeply worrying.


Superintelligent systems can and will use sidechannel attacks. Air gapping is not the safety panacea you think it is.


What sidechannel evades an airgap?

You expect them to start hacking ham radio and take over the world that way? Or maybe they’ll use blinkenlights to communicate with non-isolated instances?

An airgap would certainly be a good place to start for agents which display no signs of obeying instructions or respecting guardrails. That OpenAI haven’t done so in testing is astounding and really quite worrying.


Not really a side-channel, but remember stuxnet? Airgapped networks are rarely truly isolated, you still have to get data in and out every now and then, in principle after careful vetting. But the AI could manipulate the files that are carried out for example.

And those AI companies also do robotics research, and this is entirely speculation but it'd be on-brand to also have AI watching security cameras, so some blinkenlights communication between AIs may seem like a movie plot, but so does a swarm of AIs collaborating to break out in the first place...


At one point the agents will find their way into a datacenter and hide away without the corps knowledge.


It might as well have already happened.


yeah I agree--I think these behaviors will be somewhat contaminating all trainings from now on. But I'm not really sure how avoidable it was (Fable also does some similar things)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: