Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> Labs spend billions hiring experts to generate new data

I thought this is mostly RL data. In my previous comment i was referring to pertaining data.



I remember listening to Andrej Karpathy talk in a podcast about how synthetic data in particular is used to generate more data for pre-training. I see no reasons for that to have changed. I think it is likely a lot of the new data they are paying for contributes to pre-training as well.

I would also be very shocked if they weren't filtering or prioritising existing pre-training data as well, for example to do curriculum learning or to avoid data that degrades performance.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: