I remember listening to Andrej Karpathy talk in a podcast about how synthetic data in particular is used to generate more data for pre-training. I see no reasons for that to have changed. I think it is likely a lot of the new data they are paying for contributes to pre-training as well.
I would also be very shocked if they weren't filtering or prioritising existing pre-training data as well, for example to do curriculum learning or to avoid data that degrades performance.
I thought this is mostly RL data. In my previous comment i was referring to pertaining data.