Baking the base models on to ROM makes a lot of economic sense. SRAM for the KV cache & fine-tunes, not so much. Sure you’d get incredible speeds but it’s not scalable from a die-size or cost perspective.
Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.
I think the lifecycle for these chips could stretch far longer. If you're offering these models on a two year lifecycle, then you'd be able to stand up your top tier (wouldn't need to be frontier) at high speed. Run (for example) Kimi K3 on it and give it a brand name:
AcmeAI Carbon
Market it as your premier (only) model at high throughput. Two years later you stand up MSICs for the new state of the art with entirely new hardware, your lineup becomes:
If you just kept pushing the same model down your pricing tier over time you could still extract a lot of value from an old model, even years after it's been set in stone. Working on brand new code/frameworks? Pay to use the newest model. Working on legacy code? Use the lower tier models that will already know your legacy frameworks, pay far less and still get massive throughput. I've worked on a lot of government projects that this would be absolutely brilliant for.
The other side of this is that agent harnesses are NOT set in stone, so even a legacy model with a knowledge cut-off that's years out of date can likely still be helped quite a bit by harness and fetch behaviours that are still developing rapidly. Especially at this kind of throughput.
I'm somewhat doubtful that we will be seeing something as large as Kimi K3 in silicon any time soon.
This tech can definitely scale up from the current 8B prototype, but - at least as far as my limited understanding of the tech involved goes - you cannot just ASIC a trillion weights model due to physical size constraints.
___
Specification HC1
Model Llama 3.1 8B (hardwired)
Process TSMC 6nm
Die size 815mm²
___
So the current prototype already pushes the limits of what we can fit on a single die, and that is already likely going to limit your yield.
This is an architectural limitation that may be overcome by how you bake the MoE (mixture-of-experts) onto silicon.
If you could manage a per-die expert somehow and keep the expert routing gate relatively fast (through an interposer interconnect or doing wafer-scale Cerebras type shit) you don't need to keep the whole thing on the same die. Small dies with one expert per die on an interposer, and a very tiny router might be sufficient.
Kimi K3 is huge, though. Deepseek V4 Flash is a much more moderate model (284B total), and it works extremely well. Models of that size, and smaller, are just going to keep getting better and better. Presumably there's a threshold below which models are not generally useful or competitive, but if models-on-silicon can scale up to just 256B, that would be really remarkable.
- Deepseek V4 Flash is impressively capable. Sonnet still beats it out by a thin margin, but the real kicker is that a typical session with Sonnet at current API costs is ~$2. The same session with Deepseek is 2 cents (ha). Its even allowed me to consider offering free-with-limits API usage on my own app.
- Taalas (or competitors) have a lot going for them. If anything I feel like they need to join hands with these smaller model makers and converge in 2028
AMD can let the SRAM be on a different chip. Maybe even something similar to their 3D cache. that could increase density to 20B[1]. They could also move from 6nm to 2nm. that would probably increase density by another 3x to 60B.
Add a bunch of chips together, and you get to a server that can run a 800B model, very fast and probably significantly cheaper than others.
Right now the models are doubling in performance (by the METR time horizon metric at least) every 4 months, so 3 doublings in a year; conversely, I hear (not my field) it takes around a year to make a prototype IC and another year to turn that into mass production, i.e. if the next (late-2026 model) iPhone has a chip like this, it will likely be with, at best, a late-2024 set of weights. I think you can get open-weights models today that have performance equivalent to the SOTA-late-2024 while fitting in the RAM of a (high end) 2025-26 phone.
At some point the music will stop on training bigger models, and when that happens it will make sense to have ROM weights (or 100% analog circuits given how noise-resistant LLMs are), but we'll know when that is because the investment bubble funding the training of new models will have burst.
At some point it’s got to be good enough for the normal “phone stuff” that appeal to most users. So they wouldn’t suffer from FOMO because they didn’t wait for the next model. Every phone gimmick went through the same evolution curve until it passed the “good enough” point and eventually plateaued.
Yes, but irrelevant. While these models are improving at the present rate, the manufacturer can save money at no loss of feature-bullet-point-on-website by letting you download a model after you bought the thing and running it on normal hardware.
The rate of change to the models has to be slower than the hardware roll-out to be worth a hardware solution. If "good enough" happens before then, that just means the user gets a software solution.
You might be right but hard to tell without analyzing costs and benefits. Is a cutting edge model for phone stuff worth the slower performance and battery drain for example?
The rate of change by itself doesn’t tell you the whole story because of costs and diminishing returns. So what if your model is twice as good if it’s 10x the cost and it saves you 1ms? Everything else about phones reached “good enough for a phone” levels in years, and then got minimal generational improvements.
Transistors can be used to amplify signals, they are not limited to acting as binary switches. If you use analog rather than digital, using transistors in this way means you can replace however many transistors it would have taken for multiplying two n-bit numbers with just one; I understand capacitors can be used for accumulation, but don't know how many additional components that needs as I'm an electronics noob.
The reason we don't do this in general (any more) is that for long chains between input and output it has been much too difficult to avoid accumulation of errors. LLMs happen to be extremely resilient to errors like this, which is also why we can use e.g. 4-bit weights.
I don’t think average user _needs_ to solve frontier challenges. ”Call to Jane”, ”turn on the lights” and ”what’s the weather this afternoon” is more like it I would guess.
Ofc if the model has some critical bugs that’s another matter.
Maybe baking in a model that is "certified" to have some unconditioned truths + rest is pulled from external models/store could make sense. But AFAIK that doesn't exist and I'm not sure it can possibly be made. Perhaps society as a whole at least can work on an open corpus of training data, but I'm not holding my breath on this.
>Your examples worked on phones for over a decade.
Nope. And not only not a decade ago, right now.
If you have an Android or iPhone, you can give it clear and easy to understand instructions that Gemma 4 could complete[1] if it had tool calls on it, and that 100.00% of Claude, ChatGPT, Grok, Kimi, you name it, could understand and all complete if they had the access.
The phones will fail to complete it. I just tried Siri. I said "hey Siri", waited for Siri to come up, and then I asked one of the exact sentences you replied to: "what's the weather this afternoon?" It thought for around 20 seconds, and said "Something went wrong. Please try again."[2]
I have Wifi, I have mobile Internet, I have free storage space, I have up to date software. What went wrong is that phones have never properly connected agents, not ten years ago, not last year, not this year, and probably not next year.
But don't settle for what Google could do in 1999 by hotlinking the keyword "weather" in any query to the weather being shown in the results.
Tell your phone (any phone): "Please call back the last number that called me that is not an unlisted number, regardless of who it came from."
0 out of any phone will complete that today, tomorrow, a year from now, five years from now, ever, because phone makers are not going to let them do that.
Meanwhile, 100% of all frontier agents could complete it if they had tool calls on the phone. Which they don't, and won't ever, thanks to the duopoly.
Okay, that's a bit dismissive, I would love to be wrong!
[1] after any voice recognition to text - which does work really well on both Android and iPhone!
[2] screenshot: https://ibb.co/21rtDnfV
In my experience, they had a lot of stuff working well in the first few years they rolled out the home voice assistants - Alexa, google home, etc. But for whatever reason, they've spent the last eight(?) years silently breaking things that used to work. Stuff like audiobook playing, music alarms, or even messaging people.
Once they started seeing useful (if niche) functionality as a cost center, there wasn't really a world in which these could usefully exist. Their big bet now seems to be that LLMs will lead them to profitability - but whether that's from increased data harvesting, cheaper integrations, or because it'll be useful enough to charge subscription fees, I couldn't tell you.
Can you say this to it: "Hey Siri [wait for it to come up] - please send me an email with the temperature right now so I have it for my records." and see if it can complete the task without any backtalk or misunderstanding, and if you get exactly what you asked for. (It's a really clear request.) Should be 1 statement, no clarification, conversation, random search results, ("Here's what I found!"), etc.
A normal frontier model can do that - or Siri can do it if it is properly connected to Claude, ChatGPT, Gemini, Grok, or any other frontier AI - but previously it was never properly connected.
If it can do this task, I might have to look into this again. It counts as a success if it sends yourself any email with the current temperature and you actually get it (it can include whatever other text in the email), and a failure if it talks back, says "here's what I found", says it can't, asks you any question, sends you an email that doesn't actually contain the current temperature, just reads you the temperature and then asks if you want it to send an email, etc. Should be 1 shot.
No, they don't work. Just asked Siri the other day "what's the weather tomorrow in $LOCATION" (where $LOCATION is a broader zone and not strictly a city) and the answer was the weather in a street called "$LOCATION Avenue" in a city 150km away.
Works, sometimes. But they can fail spectacularly and unexpectedly even on very basic questions/instructions, like so simple that a hand-coded word-matching style logic could get them right 20 years ago.
And the customers can wait for the new phone released next year. These are edge models - the average customer doesn’t need the latest frontier model. Just needs to be good enough for the features you promised.
Of course, 1T SRAM isn't really SRAM, but my understanding is it doesn't require external refresh like eDRAM, is a bit easier to fab on-die than eDRAM, and is half the mm2 per Megabit compared to real SRAM (15% more die size than eDRAM)...
We also have ReRAM (Analog Computing), which also holds a promising future given its efficiency and low power. Though ReRAM of larger size is still a research area.
It's not even just that. If you just built the rom chips separately and swapped them for the RAM of a normal accelerator, it would not help at all.
The trick is that every compute element in their system has it's own small pool of ROM, instead of putting all the ram behind a common pipe. ROM is just used because it's the densest kind of memory that can be fabricated on the same process as their logic.
Yes, each rom bit can be a transistor or even a diode with a decoder circuit. Simplest Dram cell is capacitor+transistor - and you need a clock, refresh circuit etc.
Someday, I imagine model weights could even be encoded as analog resistors (memristors or similar) for even greater density
Rather base model on ROM + KV cache on DRAM is much more scalable. Also this would work great for edge devices that have a 2-5 year lifecycle.