Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There's some cool research that looks at how strongly the weights are aligned through training vs adherence to the system prompt. Like when you know a model is lying through censorship: https://arxiv.org/html/2603.05494v2

Presumably if negative guidance is in the system prompt, there's a good chance that the model would happily comply if it wasn't there.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: