Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Maybe, but that's sort of begging the question that those open weight models aren't significantly trained using "distillation"[0]

[0] not technically distillation. https://thomasdullien.github.io/posts/2026-06-15-rl-economic...



distillation is a minor piece of training data, you have to have a good foundation for it to be helpful, and even if you have good traces, you need a good RL reward scheme at the point it is used (very challenging)




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: