This isn’t a theoretical or technical problem; in fact, it’s not about the models at all: it’s about the available services getting worse. When they preview a new model and people are benchmarking it, they afford them a shitload of compute and they work great. Then, they get worse. That extra compute is surely in the billions of dollars OpenAI allotted to marketing, and especially in higher-volume periods, people talk about Anthropic‘s responses being lower quality all the time. The finances don’t indicate these companies’ current MOs are sustainable. Since self-hosted models don’t stand a chance of competing with frontier model capabilities, there is a very real chance that the available services will get much worse, and where the rubber hits the road, that’s really all that matters in the foreseeable future.