The insane valuations for selling a dream are what make VC worthwhile.
TSLA would be worth crap if it were a private company giving off dividends. It is really truly about the insane valuations driven by collective delusion.
Because you still need a good sales and product team. Speak with potential customers, understand their problems. Building the software was never the hard part.
I would maybe argue that Einstein was the most LLM-like of great thinkers.
A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.
A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.
That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?
The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?
Then it’s an expert system.
Stephen Hawking wasn’t very good at folding clothes.
The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.
I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.
Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.
You only have weights (large immutable memory), or context (small mutable memory).
Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.
Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.
For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.
So if we take a huge with enough compute (CPUs, b200s, petabytes of SSDs), we install on it both the Astra, and the toolsuite to incorporate new sensory inputs (threads/sessions), camera, microphone, temp sensors, the lot, into a new version of the model. This model is then swapped for the old model, or traffic slowly brought over, or even adjusting weights in place.
Then my hypothesis is that thing as a whole could achieve AGI.
This feels like a very close approximation on how we humans evolve our brain. By encountering new experiences/sensations, classifying them as negative or positive to us, filling it away in neurons. Or by training motor skills etc. In the end we get more connections between neurons in our brain and we are capable of more.
Bingo, LLM architecture just does not lend itself to becoming AGI. They can get really good, sure, but they will always struggle with novel input and scenarios.
The more training data that is shoved in to them, the more they'll seem to solve novel situations, but in reality it'll be things that exist in the training data.
Aka Star Trek hologram characters aren't sentient, and actually anyone who things droids in Star Wars can think of a weirdo. C3-PO just kept running out of context and trying to revert to it's system prompt.
> Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
If a model can't learn on their own to play some new game just as well as humans do, it's not AGI.
It's okay if they would take some hours or days of learning (like humans might), but if they can't do it at all during their normal operation, that's not general intelligence
> You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
But humans have general intelligence. AGI is about matching human ability, and we know this is possible in principle because brains exist
And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.
Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.
I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.
I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].
Can definitely write college level essays and have for a while. The jobs is that when LLMs first started getting popular, but aren’t quite common professors were that some of the worst students in class started writing the best essays. Now everyone complains because they can detect the slop, but most human writing is so bad. But the really good human writing is still much better.
> It's quite shocking to me how many experienced, tech-savvy people, who used to care about cookies and ad tracking - are now willingly sending their business strategies, highly confidential contracts, and intimate personal issues to a cloud provider because "it is only $0.0x per million tokens!".
Because there are more privacy guarantees there, depending on the provider. "But what if they violate their contract!" is some pretty tin-foil hat stuff.
How is this any different than a business running their website out of the cloud, assuming you are using a provider with appropriate contractual terms?
You can care about tracking and ads but still be comfortable storing your backups in the cloud, and many have been for quite awhile, even sometimes without encryption - that is totally different than e.g. Meta actively trying to track you and understand your relationship graph and your purchases etc.
I'm not sure about the tin-foil-hattedness of worrying about them violating their contract. But that's by-the-by. It is definitely not tin-foil-hat to worry about the data being taken in a breach.
Depends on how you use it. You can use AWS in a way that protects your data even if Amazon is compromised. You can even run LLM inference that way, but it doesn't seem at all common.
> They also have enough power to negotiate contracts with strong privacy provisions.
What privacy provisions would you want to add to AWS? Most of the reasonable strong privacy provisions you'd want are already there and/or available if you want to sign up for it, even including US Govt Top Secret data if you meet some approval.
If they decide to ignore the contract and violate the privacy provisions, the US government can punish them, can the company from future contracts, throw executive in jail.
If they violate their privacy agreement with me there's effectively no punishment I can get that they will actually feel. There would have to be a class action law suit and they always just settle those for some small amount and admit to no wrong doing
> If they decide to ignore the contract and violate the privacy provisions, the US government can punish them, can the company from future contracts, throw executive in jail.
To be honest, they'll just send a few million and some gold plated trinkets to the right people and it'll go away. You know that's how things work now in the US. Even if you're an actual convicted drug dealer you can have your mother say a few nice things and whoop there's your get out of jail free card. This stuff is peanuts compared to that.
If not they'll just stall in court and settle for peanuts after everyone is exhausted battling it out. That's what they've always done before the current government made it even easier. Fines are not a penalty, just a business expense.
Nobody important is going to do any time for this. Not in this America.
I am unsure if it is incapable, but it sounds hard.
They have tons of cores with 64k of SRAM each, and relatively slow paths in/out.
On a GPU you can leave it in SRAM. On Cerebras, you have to send it out of the system which is a giant bottleneck.
reply