Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Ouch. I get ~45t/s with Qwen3.6 35B A3B, and around ~70-80 with my current model Ornith 1.5 35B A3B. Local models work a treat IMO if you've got decent hardware for it.


They’re very nice for some things.

One challenge I run into is I run many agents at once, so local resource are a tiny fraction of the total inference we’re using. Every developer with his own Pro 20x, Max, etc. accounts is letting us hit literally trillions of tokens a month.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: