Selling access to LLMs via remote APIs is the “stage plays on the radio” stage of technological development. It makes no actual sense; it’s just what the business people are accustomed to. It’s not going to last very long. So much more value will be unlocked by running them on device. People are going to look back at this stage and laugh, like paying $5/month to a cellphone carrier for Snake on a feature phone.
Web apps:
- Need data persistence. Distributed databases are really hard to do.
- Often have network effects where the size of the network causes natural monopoly feedback loops.
None of that applies to LLMs.
- Making one LLM is hard work and expensive. But once one exists you can use it to make more relatively cheaply by generating training data. And fine tuning is more reliable than one shot learning.
- Someone has to pay the price of computation power. It’s in the interest of companies to make consumers pay for it up front in the form of a device.
- Being local lets you respond faster and with access to more user contextual data.
This is sort of like saying the world wide web is a fad. Many people made that argument, but a lot of desktop apps got replaced by websites even though they were supposedly inferior.
ChatGPT works fine as a website and you don’t need to buy a new computer to run it. You can access your chat history from any device. For many purposes, the only real downside is the subscription fee.
If LLM’s become cheaper to run, websites will be cheaper to run, and there will be lower-cost competition. Maybe even cheap enough to give away for free and make money from advertising?
This doesn't seem technically feasible to me. The state of the art will for a long time require a lot more hardware to run than it's available on a consumer device.
Beyond which, inference also benefits from parallelization, not just training, so being able to batch requests is a benefit, and more likely when access is offered via an API.
Well that's the problem though, those models don't come any close to being useful at all. At least not yet. And they also run much slower.
As compute increases in general, there will be larger and more capable state of the art models and it'll make more sense to just use those instead of trying to run some local one that won't give you any useful answers. Data centers will always have a few orders of magnitude more horsepower than your average laptop, even with some kind of inference accelerator card.
Do people use it for anything practical? Making stock photos maybe? I haven't really had a proper use case for it and all the random things I tried to make with it weren't good enough to use with anything. Could be useful for making concepts for real artists, but last I heard they were all too busy boycotting it.
> I haven't really had a proper use case for it and all the random things I tried to make with it weren't good enough to use with anything.
Sounds a lot like most of my early programming experiments…
Though I’ve heard on good authority that the early programmers looked past being able to calculate ballistic charts and have done some interesting things with these “computer” things.
Trying out some prompts, maybe last I used SD my mistake was going with a lower resolution to speed up generation. I literally cannot get this one to make anything that isn't a weird blob at 256px and lower, but at 512px it works fine? Weird that it's so resolution dependant. I guess some proper stuff can be made at 1024px and above.
This technology will be embedded into every OS within 2 years. People don't generally need a "super" model like GPT3/4. It will be perfectly acceptable and common to have the model change context, sync with whatever model/training data is necessary to be an expert in that context only, and associated contexts..., and prompt it in a specific domain. Client devices and internet connections are fast enough to do this in near real time today. The platforms to do all of this are being built right now by every company that creates software otherwise they will fail within 5 years.
I can already run Vicuna(llama) 7B on my 2020, 14" PC laptop at ~3.5 tokens/sec, and more speed can definitely be squeezed out.
Most future laptops and phones will ship with NPUs next to the CPU silicon. Once they get enabled in software, that means a 16GB machine can run a 13B model, or a 7B model with room for other heavy apps.
As for the benefits of batching and centralization, that is true, but its somewhat countered by the high cost of server accelerators and the high profit margins of cloud services.
Setting the M series aside, the AMD 7000 laptops already have reasonably fast memory. Faster than some old GPUs.
And that trend is accelerating. The latest rumor is that Intel is bringing back the eDRAM cache next (which means it was in planning long before the generative ai craze), and more stacked/on package memory is just around the corner.
While 7000U laptops have yet to be benchmarked, dual-channel DDR5/quad-channel LPDDR5 systems top out at about 60GB/s. (The M1/M2 by comparison is a 100GB/s, and doubles for Pro, Ultra, and Max up to 800GB/s). As a point of reference, top end consumer GPUs like the RTX 4090 are at about 1000GB/s.
My understanding is things like V-Cache, eDRAM have limited benefits for dense transformers, as they need to cycle through all/most of the parameters when running.
I don't think it's going to happen in the next few years
the prices are gonna drop like hell, but ain't no way we run models meant to run on 8 nvidia A100 on our smartphones in the next 5 years
just like you don't store the entirety of spotify on your iphone, you're not gonna run any decent LLM on phones any time soon(and I don't consider any of the small Llamas to be decent)
This is the reason why they're not going to move on device anytime soon. You can use compression techniques, sure, but you're not going to get anywhere near the level of performance of GPT-4 at a size that can fit on most consumer devices
I think we’ll see completely new architectures dominate in the near future, ousting the transformer. I am strongly suspicious that, while impressive, transformers use several orders of magnitude more compute than is “needed” for the tasks they perform—if for no other reason because the human brain performs similarly and it only draws 20 watts! And it isn’t even an engineered system, jus the product of a very, very long history of natural selection! I fully anticipate that we’ll see AI in the near future that achieves human-level performance on sub-human power budgets like the ones you’d be constrained by on a phone :)
"neat future" is very ambiguous. At the moment there is nothing even close to transformers in terms of performance. I suspect you are right in general but I'm not sure about the "near future" part, there needs to be a pretty significant paradigm shift for that to happen (which is possible, of course, I just don't see any hints of it yet).
RWKV is an attention-free architecture that's showing promising scaling at a similar level to Transformers right now! There's also recently been Hyena, which uses a new mechanism that's kind of a weird mix of attention, convolution, and implicit modelling all at once. It's shown promise as well. Remains to be seen if these competing methods will truly scale as well as Transformers, but I've got my fingers crossed. Only a matter of time!
I agree that "near future" is quite ambiguous though. If I were to disambiguate my claims, I think I'd personally expect a Transformer-killing architecture to arise in the next 4-5 years.
the only thing I can say to this is that Apple have seemed laser focused on tuning their silicon for ML crunching, that that focus is clearly now going to be amped up further still, and that in tandem the software itself will be tuned to Apple silicon.
GPUs on the other hand are pretty general purpose. And 5 years on a focused superlinear ramp up is a long time, lots can happen. I am not saying it's 100%, or even 80% likely. It'll be super impressive if it happens, but I see it as well within the realms of reason.
Apple's new M2 Max has a neural engine which can do 15 trillion flops. Nvidias's A100 chip (released almost 3 years ago) can do 315 trillion flops. Apple is not going to close this 20x gap in a few years.
FTFY, remember it takes 8 of those to even load the thing. And when the average laptop has that much compute, GPT 4 will seem like Cleverbot in comparison to the state of the art.
I think the tuning the models to the hardware piece is important, and of course there is much more incentive to do this for Apple than nvidia because of the distribution and ecosystem advantages Apple have.
But also, I don't know... let's see what the curve looks like! It's only been a couple of years of these neural engines. Let's see how many flops M3 can hit this year. And then m4 the next. Again, 5 years is a long time actually when real improvement is happening. I am optimistic.
That doesn't sound likely with the current architectures. There may be some kind of specialisation, but NN is like the chip design nightmare. We can't do chips that that many crossed lines. It's going to have to keep the storage+execution engine pattern unless we have done breakthroughs.
Well, we'll see what the future manufacturing brings, but right now we're not even at thousands of layers (as far as I know... please link if there's been more), and we'd need to be in hundreds of thousands range. Given the rate of defects also adding up and the need for some way to dissipate the heat... (almost all of that chip will be engaged while running - no chance for balancing power between systems) Yeah, still lots of challenges there.
(I'm assuming the original comment meant literally putting the network as is in the purpose designed chip)
The M2 and the 4090 are both very general purpose. In fact, the 4090 allocates proportionally more silicon area to the tensor cores than Apple allocates to the neural engine.
The M series is basically the only "big" SoC with a functional, flexible NPU and big GPU right now, which is why it seems so good at ML. But you can bet actual ML focused designs are in the pipe.
I don't think so. M chips just happen to have a really good memory subsystem and good simd performance through accelerate, so the CPU performance is pretty good.
Some stable diffusion implementations can use the NPU or GPU, or (experimentally and unsucessfully) both.
Curious, why do you think that? My knowledge is limited to marketing material and my M2 vs my 3090, and my conclusion so far would be that’s in every hardware makers marketing claims the past couple years.
> but ain't no way we run models meant to run on 8 nvidia A100 on our smartphones in the next 5 years
When I leaned about neutral networks, the general advice at the time was "you'll only need one hidden layer, with somewhere between the number of your input and output neurons". While that was more than 5 years ago, my point is - both the approach and the architecture changes over time. I would not bet on what we won't have in 5 years.
An A100 is about the size of a brick, there is no way we're fitting those 8 bricks in a phone in the next five years, without even thinking about heat management
An A100 HGX server is ~6kW of power consumption (and associated heat), while an iPhone is O(1W). I agree that a 6000x increase in energy density or 6000x decrease in power consumption is unlikely in this decade.
The human brain is also three-dimensional, heavily interconnected, and has built-in thermal management at every scale. Chips are much faster, but still operate on the essentially linear memory cells, and this limits how many matmuls you can do per second. If we can figure out true connectivity without doing tons of matmuls, then we should be able to massively cut computational demands of models.
I agree - I think for security and privacy we need it to be on-device (either that or there needs to be end to end encryption with gaurantees that data won't be captured for training). There are tons of useful applications that require sensitive personal information (or confidential business information) to be passed in prompts - that becomes a non issue if you can run it on device.
I think there will be a lot of incentive to figure out how to make these models more efficient. Up until now, there's been no incentive for the OpenAI's and the Googles of the world to make the models efficient enough to run on consumer hardware. But once we have open models and weights there will be tons of people trying to get them running on consumer hardware.
I imagine something like an AI specific processor card that just runs LLMs and costs < $3000 could be a new hardware category in the next few years (personally I would pay for that). Or, if apple were to start offering a GPT3.5+ level LLM built in that runs well on M2 or M3 macs that would be strong competition and a pretty big blow against the other tech companies.
That hardware's gonna look a lot like ASIC Bitcoin miners if an architecture to replace LLMs is popularized. General-enough purpose computing ain't going away for a long time.
I'd suspect it will actually accelerate moving everything into the cloud.
If your entire business is in the cloud, you can give an AI access to everything with a single sign or some passwords. If half is on the cloud and half is local, that's very annoying to have all in-context for your AI assistant. And there's no way we're getting everything locally stored again at this point!
Right, this is why StabilityAI is getting in bed with Amazon, so private, fine-tuned models can operate on all your data sitting out there in S3 buckets or whatever.
What's been so interesting with the explosion of this has been how prominently the corporately-driven restrictions have been highlighted in news and such.
People are getting a good look in very easy to understand terms at the foundational stage at how limiting the future is to have this just be another big tech controlled thing.
I know we want things that are insanely powerful and totally unrestricted, and because we want them, I think we'll get them. And then I genuinely think this tech is going to end in tears.
They have said that the alignment actually hurts the performance of the models. Plus for creative applications like video games or novels, you need an unaligned model otherwise it just produces "helpful" and nice characters.
The character simulacrum used by an LLM tends to be the result of "system" prompts that set by the service you are using. GPT-N isn't exactly trained to be helpful and nice, but ChatGPT has system prompts describing the character it should be performing as. If you work with just GPT-4, you can get more zany outputs.
That said, OpenAI does use RLHF, which does bias the model away from raw internet madness and something that OpenAI wanted at the time of training. A lot of models haven't gone through rigorous RLHF, though.
As a side note, RLHF might be the best alignment technique we currently have in practice, but it is not decisive. It has been noted in multiple experiments that RLHF can just train a model in how to trick the human reviewer, if tricking is easier in practice than doing a think the human review wanted. So this isn't even really seen as aligning a model by alignment researchers. At least not an approach that can scale with the increasingly intelligence AI models.
Alignment is an unsolved problem. None of the current stronger models are "aligned", just tuned in ways that weight some biases more than others, but even that is dependant of the features of their inputs.
On this topic, Apple is the sleeping giant. Sleeping tortoise maybe. Everyone else has been fast out of the gates, but Apple has effectively already been positioning to leap frog everyone after a decade+ of M1 chip design. Ever since these chips launched, the M1 chips have felt materially underutilized, particularly their GPU compute. Have to believe something big is going on behind the scenes here.
That said, wouldn't be surprised if the truth was somewhere in between cloud-deployed and locally deployed, particularly on the way up to the asymptotic tail of the model performance curve.
What would a "leap frog" look like, in your mind? I'm struggling to imagine how they're better positioned than the competition, especially after llama.cpp showed us that inference acceleration works with everything from AVX2 to ARM NEON. Compared to Nvidia (or even Microsoft and ONNX/OpenAI), Apple is somewhat empty-handed here. They're not out of the game, but I genuinely see no path for them to dominate "everyone".
My guess is a leapfrog would have more to do with how LLMs are integrated into an operating system, rather than just coming out with a better model. I don’t think we’re gonna get a substantially more capable LLM than GPT-4 anytime soon, but fine-tuning it to sit on top of the core of an operating system could yield results.
Feels like Microsoft already beat them to the punch. Their ONNX toolkit has better ARM optimization than Apple's own Pytorch patches, and their collaboration with OpenAI places them pretty far ahead of the research curve. I'm convinced Microsoft could out-maneuver Apple on local or remote AI functionality, if they wanted to.
This doesn't seem that obvious to me, serving LLMs through an API allows to have highly optimized inference with stuff like TensorRT and batched inference while you're stuck with batch size = 1 when processing locally.
LLMs doesn't even require full real-time inference, there are applications like VR or camera stuff where you need real-time <10ms inference, but for any application of LLMs 200-500ms is more than fine
For the users, running LLMs locally means more battery usage and significant RAM usage. The only true advantage is privacy but this isn't a selling point for most people
You're still thinking in terms of what APIs would be used for, rather than what local computation enables.
For example, I'd like an AI to read everything I have on screen, so that I can ask at any time "why is that? Explain!" without having to copy paste the data and provide the whole context to a Google-like app.
But without privacy guarantee (and I mean technical one, not a pinky promise to be broken when VC funding runs out) there's no way I'd feed everything into an AI.
We are very close to optimized ML frameworks on consumer hardware.
And TBH most modern devices have way more RAM than they need, and go to great lengths to just find stuff to do with it. Hardware companies also very much like the idea of a heavy consumer applications.
That's what pruning is, but it's not that straight forward and has limits. Finetuning a smaller model on the output of a larger one is much more flexible and reliable.
GPT 3.5 is probably a 13B Curie finetuned on the output of full size GPT-3 175B, to give you an idea of the technique.
That is smaller than the third smallest StableLM and the same size as LLaMA-13B which can run at useful speeds off of a smart phone CPU.
GPT-3.5 is much worse at "complex" cognitive tasks than Davinci (175B), which seem to indicate that it's a smaller model. It's also much faster than Davinci and costs the same as Curie via the API.
It's clearly a smaller model, but I'm very skeptical that it is 13B. It is much more lucid than any 13B model out in the wild. I find it much more likely that they used additional tricks to scale down hardware requirements and thereby bring the price down so much (int4 quantization, perhaps? that alone would mean 4x less hardware utilization for the same query, if they were using float16 for older models, which they probably were)
I'm sure they're tweaking lots of things under the hood, especially now that they have 100M+ users. It could be bigger (30B?, maybe 65B) as coming down from 175B gives quite a lot of room, but the cognitive drop from Davinci gives away that's it's much smaller.
People fine-tuning LLaMa models on arguably not that much/not the highest quality data are already seeing pretty good improvements over the base LLaMa, even at "small" sizes (7B/13B). I assume OpenAI has access to much higher quality data to fine-tune with and in much higher quantity too.
I have been playing with all the local LLaMA models, and in my experience, the gains that are touted are often very misleading (e.g. people claiming that 13B can be as good as ChatGPT-3.5; it is absolutely not) and/or refer to synthetic testing that doesn't seem to translate well to actual use. Using GPT to generate training data for fine-tuning seems to produce the best results, but even so, GPT4-x-Alpaca 30B is still clearly inferior to the real thing. In general, the gap between 13B and 30B for any LLaMA-derived model is pretty big, and I've yet to see any fine-tuned model at 13B work better than plain llama-30b in actual use.
So I think that 65B may be a realistic estimate here assuming that OpenAI does indeed have some secret sauce for training that's substantially better, but below that I'm very skeptical (but still hope I'm wrong - I'd love to have GPT-3.5 level of performance running locally!).
Agreed, there is way too much hype about the actual capabilities of the LLaMa models. However, instruction tuning alone makes Alpaca much more usable than the the base model and to be fair even some versions of the "tiny" 7B can do small talk relatively well.
> Using GPT to generate training data for fine-tuning seems to produce the best results, but even so, GPT4-x-Alpaca 30B is still clearly inferior to the real thing.
Distillation is interesting and it does seems to make the models adopt ChatGPT's style but I'm dubious that making LLMs generate entire datasets or copy/pasting ShareGPT is going to give you that great of a dataset. The whole point of RLHF is getting the human feedback to make the model better. OpenAI's dataset/RLHF work seems to be working wonders for them and will continue to give them a huge advantage (especially now that they're getting hundred of millions of conversations of people doing all sorts of things with ChatGPT)
I think it may be naive that people believe that the deciding factor on how these things are used is likely to be "chip speed." or "efficiency on the machine."
I wish we were in that world; but it more likely seems like it would be "Which company jumps ahead quickest to get mindshare on a popular AI related thing, and then is able to ride scale to dominate the space?"
REALLY hope I end up being wrong here; the fact that so many models are already out there does give me some hope.
I don't that's true in the context of businesses because they won't want their data to be leaked and/or used for other clients. The more data from your company you can feed the AI, the more productive it will be for you. I'm not just talking about semi-public documentation, but also things like emails, meeting transcript, internal tools APIs, employee details, etc.
If the AI service provider uses your data to help better train their AI, it will be blacklisted by most companies. If you keep them in silos, the centralisation will offer almost no benefit while still being a very high privacy risk. The only benefit they get is that it allows them to demo it and see it's potential, but no serious business will adopt it unless you also provide a self-hosted solution.
I think the only people who will truly benefit from using cloud services as a long term solution are personal users and companies too small to afford the initial cost of the hardware.
That seems hard to believe for businesses which already rely on Office, Teams and Sharepoint, since Microsoft will be making its version of ChatGPT available for all its products, and the integration will be too hard to pass up on.
Microsoft is in a different situation because everyone is already forced to trust them with their OS and o365. For better or for worse, there are no current alternatives to Windows and the office suite for most businesses. If you already login to your OS with a Microsoft account and process your data in Excel, adding an AI tool on top of it is not a big jump. Very few others are in this situation.
For every other AI service providers, good fucking luck getting clients to trust you. I expect we will see a lot AI services that offer a cheap and easy to use cloud AI subsidized by a very expensive self-hosted version. I also expect a lot of data leaks and many high profile incidents where an AI creates a document or code that includes sensitive data from someone else (hard coded passwords, API keys, etc.).
Even for a large company like Autodesk or Adobe, you might trust them with your engineering drawings and your new product design, but would you feel comfortable uploading your code base for internal tools, employee files, email communications, etc. to them? It's gonna be a hard no for a lot of businesses
Having more users helps with reinforcement learning, but as a user, I want an unaligned AI that isn’t constantly babysitting me with bullshit about what it can and cannot do, so there’s like a negative network effect, lol.
There will be a time when LLMs need data persistence to "improve our user experience". The LLM will act like a "friend" that will remember you when you come back.
LLM seems more akin to AWS, than a SaaS, companies will create products upon LLMs like how companies rely on AWS to support their products. The build vs buy calculus may tip heavily towards build once they can run on device with good user experience, no need to pay for cloud compute any longer.
> The build vs buy calculus may tip heavily towards build once they can run on device with good user experience
Hahahahahaha... oh wait, you're serious? Let me laugh even harder.
Have you used any commercial software in the last 25 years? Garbage web apps have replaced very nice, performant local applications across the board. My stupid fitness tracker app (that should be a 10 MB sqlite DB) instead fails to even open without an internet connection.
Is your theory that companies will suddenly decide they hate getting money and love paying money for developers to create great user experiences?
This is mostly why the future of computation only makes sense monetarily if you have everyone shift to a thin client. So, banning GPUs is likely considered a "necessary evil" by the BigTech cognoscenti for accomplishing that goal.
When radio first started, people read plays written for the stage, because that's what they knew and what they had. Later people learned to write for the medium and make radio native entertainment.
Same thing happened when TV arrived. They did live versions of the radio entertainment on a set in front of a camera.
Selling access to LLMs via remote APIs is the “stage plays on the radio” stage of technological development. It makes no actual sense; it’s just what the business people are accustomed to. It’s not going to last very long. So much more value will be unlocked by running them on device. People are going to look back at this stage and laugh, like paying $5/month to a cellphone carrier for Snake on a feature phone.
Web apps:
- Need data persistence. Distributed databases are really hard to do.
- Often have network effects where the size of the network causes natural monopoly feedback loops.
None of that applies to LLMs.
- Making one LLM is hard work and expensive. But once one exists you can use it to make more relatively cheaply by generating training data. And fine tuning is more reliable than one shot learning.
- Someone has to pay the price of computation power. It’s in the interest of companies to make consumers pay for it up front in the form of a device.
- Being local lets you respond faster and with access to more user contextual data.