Centralized inference can easily increase batch size, leading to huge efficiency gains in the usual scenario where most users have just one or very few session. Using local resources efficiently requires some way to increase the batch size. I'm not sure if we are there yet.
Also, an entire family can share one big box that can run a good model. Not an entire household, but an entire family no matter where they live in relation to the box. They could all access it over the internet, or even simply over the telephone. When it becomes another family member/servant, a price tag in the low $10s of thousands seems a lot more reasonable - and this is accessible now.
I could absolutely see local AI taking the place of voicemail/call screening completely. Call me and you get my AI, who will route the call to me if approved, or even choose to service the call itself. If a friend of mine calls who doesn't have $20K to spend on a great AI rig, they would certainly have permission to steal otherwise wasted cycles from mine. An answering machine isn't much different than an issue tracker, and some of those are 97% AIs having perfectly intelligible conversations with each other, and 3% humans being eagerly serviced by AI.
The old objection was "this is so hard to set up, nobody wants to run a server!" Now it can set itself up.
tldr; you don't have to rent from some data center. You can get high utilization out of a local box. This 1) puts a hard ceiling on what a data center can charge, and 2) they won't be able to compete on privacy, which locally can be complete.
I think we'd need something different to make that a reality today. It's likely that decentralized compute eventually wins out here as well. It usually does. But without architectural changes, it will take a couple of years before the current single-conversation flow is truly usable on consumer-grade hardware.
Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks.
The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
> The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
Agentic AI has pluses and minuses for cloud efficiency. The plus is that usage could be very bursty, but the minus is that agents will more fully utilize a local system. The main disadvantage of local AI is that you would be paying a large amount for a system mostly doing nothing. If it's constantly working on different projects and integrating that data, you get use out of every penny that you spent. Every GPU you added would instantly make the thing smarter.
What's more, your local AI could offload an agent to the cloud if it needed to. It could do this rationally, based on your personal desire for privacy.
Jevon’s paradox suggests that total datacenter resource consumption will increase. It doesn’t say that people will choose datacenters over their own personal hardware when the latter is sufficient.
Personal hardware is only sufficient today for some tasks. As data center power efficiency, and large sparse MoE task efficiency increase, personal computing will continue to lose out.
Centralization without proper controls against monopolization becomes sloth and gluttony. If american labs were constrained like china, their models would benefit.
The abstract benefits are quickly outstripped. The same way adding more highway lanes never improves gridlock. Its inducement.