Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The size of model we're talking about running doesn't need much if any dram.


The chatjimmy demo is using a model that needs 6-18GB of VRAM. That's not exactly trivial.

I could see it being feasible to get a Qwen-3.6-27b type of model done on something like this. Qwen-3.6-27b at 18tok/s would be a game changer.


right, but that's a reticle size chip. to put something in a phone it has to be ~10-30x smaller




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: