Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Inference costs will go down massively once they use the upcoming GPUs. I estimated that a model like GLM5.2 will be around 0.03USD/M output tokens in 2 years when the Feynman GPUs will be available in 2028. And this did not even consider architectural efficiency improvements. In mid 2027 we will already see a 10x reduction once everyone has switched to the Ruby architecture.

It will be feasible for everyone to have 20 different agents running at all times. A new world is coming



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: