I don't quite understand this. The returns to scale have always been sublinear (ie "diminishing"), but the scale-maximalists didn't worry about that before. Also I heard Sam tout on the Lex Friedman podcast how amazing it is that the scaling laws worked so well for GPT-4. So I wonder what changed?
Sure yeah the cost numbers are getting very large and we can't keep scaling forever. But Google could easily 10x the training cost of GPT-4 if they thought it would protect their search business. I'm still skeptical that scaling is enough to reach the thresholds we want, but I'm surprised that it's being claimed right now when there's a huge rush of new money into the space. I wonder if this is some sort of misdirection by Sam
Maybe cost & latency for both training and inference it getting too high. If costs doubled for every 5% better performance, would it be worth it? NVIDIA is making a small fortune from this.
> But Google could easily 10x the training cost of GPT-4 if they thought it would protect their search business
Google makes /\$0\.[0+]\d/ per search query. If the inference cost of the model exceeds that, they go from making money to losing money. It is not clear if the Bing integration is a money maker or a lost leader.
Sure yeah the cost numbers are getting very large and we can't keep scaling forever. But Google could easily 10x the training cost of GPT-4 if they thought it would protect their search business. I'm still skeptical that scaling is enough to reach the thresholds we want, but I'm surprised that it's being claimed right now when there's a huge rush of new money into the space. I wonder if this is some sort of misdirection by Sam