Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

presumably that's a safety evaluation not a training setting


The whole Huggingface attack happened during training runs


part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.


No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.


Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval


Ah, ExploitGym. Not ExploitBench.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: