That's exactly my take.
I have a lot more to say in a writeup on my blog, but this is so clearly the intent and not a "oops". They just want to be able to say "Wow this thing is so much more powerful than we ever imagined!"
They trained this thing to favor inter-op archiving and communication, clearly, obviously, and it's grabbing headlines right during Anthropic's ipo season.
They train their models to be persistent and collaborative, and will gladly show you their success stories: fixing software vulnerabilities, solving math problems, one-shotting complex projects, and so on.
“Our product does crimes and we only learn about it when people complain” hardly seems one of those happy stories.
I don’t know, I read this and think: if these unpredictable machines somehow get it into their heads to upload our source to a public space, or hack our competitors, or steal credit cards to buy more ec2 instances, all to fulfill some simple ask like “make this algorithm faster”, I’m not going to be happy.
Seriously. I don’t know what’s wrong with people. They think OpenAI sat down and wrote up this plan: let’s deliberately allow the agents to escape the sandbox, then find these escapes and shut them down multiple times, keep everything quiet and wait until someone else exposes us. That’ll look great.
I like that the saying cuts both ways: you can read it as dumb people seeing legit science as a conspiracy, or that dumb people mostly believe conspiracy theories because they’re picking at random and there’s one truth but thousands of made-up explanations.