Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
NooneAtAll3
19 days ago
|
parent
|
context
|
favorite
| on:
GPT-6 Astra
> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
thatguysaguy
19 days ago
|
next
[–]
presumably that's a safety evaluation not a training setting
estearum
19 days ago
|
parent
|
next
[–]
The whole Huggingface attack happened during training runs
thatguysaguy
19 days ago
|
root
|
parent
|
next
[–]
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
cubefox
19 days ago
|
root
|
parent
|
prev
|
next
[–]
No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
estearum
19 days ago
|
root
|
parent
|
next
[–]
Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
cubefox
19 days ago
|
root
|
parent
|
next
[–]
Ah, ExploitGym. Not ExploitBench.
azeemba
19 days ago
|
prev
[–]
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
...why exactly are they training for that?