Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

About knowing whether a scenario is fictional, there was an interesting finding in Anthropic's J-Lens research.

When they benchmarked the model to evaluate whether it would try to blackmail someone in a contrived scenario, the J-Lens showed "fake" and "fictional" in the workspace.

And if edited out, the model was more likely to do the blackmailing.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: