Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

People are not talking enough how huge this is for robotics and IoT. Current robotics arhitectures are limited by tok/sec. How cares if its not a Fable model?

This move undercuts NVIDIA directly.



With AI models like Mixture of Experts, many of those experts will be the real target here, as polished, refined and little to no change, they become fine candidates for being locked into silicon. Who knows, add some SRAM in there and small changes to those experts could be carried out without needing new silicon.

Maybe AI models may become reduced to a collection of tiles you add to a chips one day, maybe sooner for some areas as you say, motor control for balance, vision systems, speach recognition systems etc, broken down, for robotoics, much is already there and just cost of battery/power holding much back.


An "Expert" is really just an unfortunate name for what amounts to a dense part of a sparse matrix and that's also an oversimplification.

It doesn't actually specialise in anything in particular that one can point to.

For this reason you can really transfer them between models.


Is it still true? I'd assume you should be able to freeze the matrix and unfreeze an expert block, before feeding particularly chosen training data. Or that doesn't work?


That's more or less the idea behind Low-Rank Adaptation, or LoRA.

There's also Mixture of LoRA Experts, which instead of slicing up the model and routing through that, routes through different LoRAs.

But it all comes with tradeoffs, as you have to train and run the gating network doing the routing, which also comes at a cost.


no you're wrong. people are not talking about the load bearing seam that this strong decision has revealed towards veterinary care.

the implications for mental health of pet rats is huge.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: