OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.
Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.
I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.
Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.