Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt.

Anthropic have also observed similar things, so while it seems to me that OpenAI’s level of control is more of a dumpster fire, it’s by no means a unique issue to them.



It's almost as if it's not actually possible to align an unknowable mystery box of floats.


Good thing we're not trying to deploy them into fully autonomous weapons or anything....


I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: