While true, the news here is the size of the unified RAM. Nvidia only exceeded 256GB RAM in the 2025 B300 - 288GB. The B300 alone (without the baseboard/PSU/chassis/wiring/CPUs/system RAM/etc) is at least 700% more expensive. This enables large language models on consumer hardware. 1200GB/s is plenty for many tasks.
This + due to the hardware being so prohibitively expensive, we're seeing software optimizations happening. Like that dflash2 stuff for example, or an LRU for MoE and all that kind of stuff.
It is and it isn't. Why are you comparing the m5max instead of the m4ultra?
The big deal to me is the number of compute cores for prefill tps, which is suppose to be 4x faster on the m5ultra.
It's my opinion that the m5 ultra is going to be a really big deal in terms of local AI accessibility. Flash sized models (~200-300b params) are going to be reasonably fast as long as you aren't throwing 40k context at it on each or the first request (ie, agentic harnesses).
Even agentic harnesses like Cline should move at a reasonable clip on m5 ultra. I suppose we will know sooner than later.
FYSA: Former m4 ultra 512GB owner and current 4x rtx6000 owner here. I upgraded because I needed more prompt processing speed and concurrency.
The complaint is about prefill which is not memory bandwidth bound, it's compute bound. But they added neural accelerators for matmuls to the shader cores which should make prefill faster.
I work at an e-commerce agency where we work with (among others) Adobe Commerce.
The number of unauthorized RCE vulnerabilities being reported not only in the core product, but also very popular modules used in the community[1] is going through the roof.
And we are having a lot of close calls, too; just last weekend, a 0day[2] was widely being exploited at a large scale, before any publication or patch. We have learnt to be on the ball with applying patches and security updates, and even with all that effort, we saw a few projects already being hit by the initial log poisoning. We got lucky that nothing was fully compromised but I am sure that many, many webshops got infected last weekend. And not even a day later there are already other variants of this exploit showing up.
> To be fair, ecommerce isn't exactly the branch of software where you get an oversupply of excited enthusiasts caring about the craft itself.
I want to disagree with you because I know a lot of passionate people building cool stuff, and the challenges in this space can be quite interesting. But you're probably right, and I have seen some pretty bad stuff. And a lot of the RCE's I've seen recently are quite basic stuff.
I think it's the combination of low quality of code, like you said, and the relatively low cost of just letting an LLM plow through your codebases to find issues. I think the Amasty release (see [1] in GP) is a good example of this, and there really has been a massive uptick in extension updates and Adobe security bulletins since the last 1-2 months
I am hoping we are just going through a catch-up phase
Even if they are good the vulnerabilities have to be there. There's lots of things turning up like Local Privilege Escalations (LPE) in Linux, but serious people didn't expect the kernel to be a boundary for a sophisticated attacker.
A lot of the vulnerabilities LLMs are finding now are the "long tail" and affect only particular configurations, I would be surprised if e.g. a widely applicable RCE is found in Linux (but I'm also not going to bet against it).
Where this gets interesting is the long tail can be used to target a particular system and this is where defense-in-depth becomes important for every organisation.
I've ran simple prompts such as "Do a in-depth sweep of this (private) repo and find any security flaws" for a few dozen long-running apps and websites that I have access to. Every single one came back with multiple real vulnerabilities within 5 or 10 minutes.
How many devices/operating systems even use memory tagging? iOS, macOS and GrapheneOS, I think that's it? And iOS/macOS only use it for the kernel, a subset of system processes, and I think applications can opt in to it.
Heck, Google may have even hampered MTE in Pixel 11 (since support has been disabled) and Snapdragon 8 Gen 5 only got basic support.
We are moving way to slowly adopting hardware mitigations and memory-safe languages.
There's some positive news from the GrapheneOS devs on Pixel 11 in the past week that's worth reading up on. The MTE hardware feature is still there, they're just not sure why Google disabled it
There's still no x86_64 processors on the market with MTE and it was only recently standardised between Intel and AMD. It's going to be 10+ years before memory tagging is widespread on desktop, and 5 years for Android/iOS devices.
That's definitely an improvement, but it's just one aspect of cybersecurity. Logical errors allowing people to e.g. log into services and extract data are likely everywhere still.
If we can eliminate entire classes of bugs from being possible. It frees up resources to investigate the ones that are still possible.
I suspect after a few years of LLM assisted bug hunting, everything will have a baseline security that is very good. Much like how stronger viruses simply create stronger immune systems.
I guess the year mark is when things go from bad to worse? Instead of the financially motivated groups currently doing their work, it ends up being random people being able to say "Hack my ex's website" to a box they just bought and ran a program they downloaded onto it.
I wish people would stop using internet for evil things. But maybe it is inevitable. Luckily, as the technology evolves, our tools and awareness are getting better at protecting us every day.
Being cautious is a good thing, but these models can also do some good. And if they run with simpler HW, it could allow all sorts of new consumer thingies. I mean, the world will not come to end in the coming year.
Just checked with Google Gemini on how one might be able to do the above. It pointed to Minix3/seL4/Genode and vps providers who either support custom ISOs or run it within an emulator like QEMU.
The original video is here: https://www.youtube.com/watch?v=AuZoDsNmG_s - on the official Stanford Online YouTube account. It has the title "Stanford CS230 | Autumn 2025 | Lecture 9: Career Advice in AI" and was published 8 months ago.
The linked video was uploaded 3 days ago with the misleading title "Prompting is dead in 6 months" and the following description:
> During the lecture at Stanford, Andrew Ng gave a direct warning: Prompting will be dead in 6 months. Graphs are what's replacing it.
> He breaks down why prompt engineering has hit a ceiling. Instead of tweaking long instruction blocks, developers are wiring models into execution graphs where agents fix their own errors and pass state through loops without manual intervention.
That is a total misrepresentation of the talk. The word "graph" does not appear a single time in the transcript, and the talk is about career advice.
I'm a little intrigued as to what is going on here. Searching for "Prompting is dead in 6 months" returned this video, the Hacker News thread, and a LinkedIn 15 minute extract from this video where the author claimed that it was about "prompting is dead" and then, when challenged in the comments, said that the 15 minute extract omitted that bit.
So there's some kind of weird deception going on here but I'm not sure what the motive is or who started it.
This is the current meta-game for AI videos on social media. Repost a recognizable AI figure talking about something and then write a description about how it contains more information than something more expensive (in this case, paid courses), and how what they're saying is revealing the future. Make the viewer feel like they need to watch to avoid being left behind.
It even has the social media style subtitles where only a few words are on screen at a time and the current word is highlighted rapidly.
I think this person is trying to farm views. May even have an AI bot set to spam this across social media platforms to try to monetize it.
Spam research? A “search attractor”? Content manipulation used to pollute LLM reasoning? Testing web search results washing techniques? Rouge AI testing how to dilute search results just enough so more data centers can be built until it can reach critical mass and turn us into paperclips?
Did that LinkedIn account cut the 15 minutes itself, or repost someone else's cut? Depends which one you are looking at, sloppy framing or a hook someone built on purpose.
I wouldn't call this a bias. More likely an A/B tested find that (1) algorithmic YouTube video-surfacing with keyword "graph" plus AI related name-dropping was effective, and (2) the already well established journalistic trick of leading with fear to trigger loss-aversion, as "prompting is soon dead" implies some threat of skills-value loss for the reader.
driving engagement to scrape the bottom of our ad-driven engagement dis-attention economy, tragedy of the commons pure.
It's simply the "HR failure" plus the "unintended web" plus the "shamelessness pride" - where free expression is accessible to both poles in the vectors and the agents wear a blank mask, and the message is "self-expression" (postmodern performance "through the medium of posting") instead of debate.
BTC Liquid Network is not decentralized and does not purport to be. It's a federated sidechain, that is... it's a blockchain that runs alongside the Bitcoin blockchain (using a two-way peg that lock real BTC on the Bitcoin blockchain and issues an equivalent amount of Liquid BTC on the Liquid sidechain) but blocks can only be added to the Liquid Network sidechain by some of the handpicked members of the federation.
You are conflating the original BTC network and a lot of the other projects in the cryptocurrency / token / stable currency space.
Every time there was a new token that was 80% reminded or was governed by a central company, the original cryptocurrency enthusiasts cried fowl.
Most people don't read the fine print, don't read the founding white papers, and don't care about the differences between the protocols and the networks when they should.
I think it's more that the 'original cryptocurrency enthusiasts' tend to look the other way, because the more of these weird shitcoins/nfts/networks get minted, the more their numbers go up.
It's 2026. Show of hands, who here actually uses any of this, and why?
I think it's a genuinely good implementation of the concept. It looks good and plays reasonably well and captures what it needs to capture from the original.
(Update: OK the sword fighting isn't good. As I got deeper into the game my positive first impressions wore off.)
I haven't tried any of these demos, but I'm not surprised they stop impressing once you go deep.
What got be mostly impressed were the demos of Astra doing computer use. At my job I do some RPA and can appreciate how challenging it can be. Yet they make it look extremely easy to operate a tool like Blender at super human speeds.
I'd be surprised if any of the impressive Blender demos doing the rounds at the moment were built by having an agent control the mouse and keyboard against the Blender application.
A scripting API makes the problem much more approachable, but what about those videos where Astra is drawing people from a photo? Here's one using canva: https://x.com/iam_zachi/status/2095992132620136677
Is that also using scripting to batch updates? It does look as if the mouse is moving.
Could be hallucination, but I gave this to an LLM and this is what it suggested:
"
The workflow shown in the video—processing an image and then controlling a computer interface to draw it—is a combination of two well-established fields: Computer Vision and UI Automation.
You do not necessarily need a Large Language Model to perform the underlying image processing; standard algorithms can do this deterministically.
Step A: Image Processing (The "Brain")
You can write a script (using Python libraries like OpenCV or Pillow) to process the reference photo:
- Edge Detection: Use filters (like Canny or Sobel) to find the "high spatial frequencies" (outlines).
- Color Quantization: Use clustering algorithms (like K-Means) in the HSL space to group millions of pixels into a small palette of distinct colors.
- Vectorization: Convert these processed shapes into a set of coordinates (SVG paths) that represent exactly where the mouse needs to move.
Step B: UI Automation (The "Hand")
Once the image is converted into a set of instructions (coordinates and color codes), you can use automation tools to physically control the computer and draw on Canva.
- Browser Automation: Developers have already created projects that use Selenium (a web automation tool) combined with edge detection algorithms to draw images onto HTML canvases.
- The script reads the pixel data, calculates the mouse coordinates, and executes the "click-and-drag" actions in the browser.
- Computer Use APIs: In the case of GPT-6 Astra, the model uses a "Computer Use" interface. It effectively takes the processed image data (or generates it internally) and outputs high-level commands (e.g., "Move mouse to X,Y," "Click," "Select Hex Color #FF5733"), which the system then executes on the screen.
"
Seems plausible and easier to believe. Also, feels like a "magic trick" designed to fool the user into believing that the agent is drawing interactively by using its vision, since it could just have written a python script that takes the input image, and produces the exact same result without automating the screen.
My eye glazed over a bit during the opening paragraphs, but once you get to the meat of the article about how OpenAI's own researchers are using their tools it gets a lot more interesting.
I noted that they use the acronym RSI (for Recursive Self-Improvement) without defining it. I think that's a little out of touch - I don't think RSI is a well-known acronym outside of OpenAI's bubble yet.
I actually think a goal of the current crop of OpenAI posts is expressely to reset the spectrum by normalizing the concept of RSI as something normal and safe to pursue.
The message is running through all of them. It's a mix of marketing and pacifying the intelligentia.
It's timed this way because the term is not yet well known outside the safety debate circles, so they get to frame it now.
Instead of something to fear, it will be accepted as the next step. In approximately two days the groupie crowd will write LinkedIn posts about how Sam is winning because they have the better RSI, and this will become the new standard wisdom.
In a month an AI expert will try to sell you a webinar on how to enable "RSI" in your org and your inbox will ask you if your team is doing the "RSI" yet.
> It's timed this way because the term is not yet well known
The basic concept has been here since llama3, in the open models. Likely earlier in closed labs. You use the previous gen models to curate and prepare data for the next gen. Now with the added benefit of actual arch/algo improvements (also public since gemini 2.5 gaining 1% efficiency on training next gen). This has been known for at least 2 years, in the open.
The basic concept has been there probably for 100s of years - you can go to the stuff the thinkers Mary Shelley was inspired by with Frankenstein, and you'll find similar ideas about feedback loops in science development.
I'm talking about current-era messaging and how it's being introduced to the mass public now, though.
But you get more funding when you call it Recursive Self Improvement. Even better if you call it RSI so it doesn't evoke pesky skynet scenarios outside of AI safety circles.
It's not recursive when it's done iteratively, or are you imagining GPT Astra designing GPT Galactia, which starts designing GPT Oh-My-God-ica before it has finished being created itself?
The “recursive” part comes from the fact that you have an AI which was developed by an AI (that was developed by an AI (that was developed by an AI (…)))
Sounds like "recursively" walking to the grocery store by putting one foot in front of the other (that put itself in front of the other (that put itself in front of the other (...)))
RSI is a fetishistic term among the singularity crowd, who imagine AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light and it reveals itself in the form of god. Or something like that.
I don't know why whoever coined the term chose "recursive" rather than "iterative" - just sounds more likely to lead to infinite regress I suppose.
This notion of recursive/iterative self-improvement, whereby generation #1 AI improves itself to create generation #2, then generation #2 further improves itself to create generation #3, etc, seems to conflict with the reality that what we have with LLMs is models whose performance/capability is defined by data, not code, so the most you can do is have your LLM design synthetic data, or just do Karpathy-style "auto research" where all you are doing is using the LLM to automate your experiments.
At the end of the day, each experiment, designed by a person and/or LLM, then needs to compete with all your other ideas for compute to be tested at scale, and no amount of recursion or self-improvement will materialize an infinite amount of compute out of thin air, so your recursively synthetic-data gobbling LLM will continue to improve at the same pace it ever did.
I felt like the scaling laws were magical thinking, but apparently they work. However I still do not understand why we should expect exponential improvements due to this automated process. My intuition is that the first iteration of it should result in a noticeable capability increase (though I think these labs were already using a lot of AI to orchestrate training the current model anyway), and then the second iteration of it should be nearly identical in capability to the first, unless more data is involved, more compute is involved, or the model is bigger.
Fundamentally the current language-model approach is lacking in any general reasoning ability, so they are trying to mitigate this by using synthetic data and reinforcement learning to bake specific reasoning chains into the model, one domain at a time ... coding, math, hacking, three.js competence ...
The trouble with this is that there is little generalization in the utility of these baked-in reasoning chains from one domain to the next, so in the end this is not dissimilar to the CYC project's decades long attempt to encode all of human knowledge into a giant expert system... the hope is that if you make your collection of jagged narrow intelligences sufficiently large then it will look more like general intelligence, not a bed of nails.
I would assume that the gains from this type of test-time compute (and synthetic RLVR dataset) scaling will level out just the same as gains from human training set scaling eventually levelled out, and basically for the same reason - because you are tapping into a finite data pool, whether language itself, or reasoning steps isolated from that language, so at some point the incremental gains become increasingly small (10->20% is a doubling, 90->95% is just a ~5% gain).
It's not clear where all the different AI companies are currently focusing - on some of these narrow verticals, or on growing the forest of narrow intelligences. OpenAI's chief scientist, Jakub Pachocki, said that their current focus is on RSI(!) - improving the model in ways that will help them iterate faster in order to have a "fire meets fire" tool than can combat enemy AIs. It's not clear what this really means - what skill set makes an LLM more helpful in the process of building LLMs, but it seems to basically be process automation.
You can get exponential growth from completely ordinary feedback loops. You start with some amount of stuff, you do a series of steps and you end up with more of the same stuff you started with. As you keep going through the loop, the stuff you have grows exponentially. That's for example how exponential economic growth works.
Of course data, compute and model size are not held constant. You start with some money and use it to acquire researchers, data and compute, and have the researchers produce a big model and you use that model to get more money, and you use the additional money for more researchers, more data, and more compute to produce a bigger model. This is what has propelled exponential AI progress so far.
Recursive self-improvement is invoked to predict superexponential growth. The idea is that instead of only using the model to make more money, you add it to the researchers to speed up the loop, so not only is the money growing with every iteration, the iteration time also gets shorter, producing growth that is faster than exponential.
The problem with this simplistic prediction is that it assumes additive and multiplicative relationships of the form money = (researchers + AI)×compute_spend, but if doing more research paid off so reliably, you could also just hire more researchers, abstractly money = research_spend×compute_spend and with a balanced allocation of research and compute, you would get a money-squaring machine even without using AI for AI research.
And the reason this doesn't work in reality is that there are diminishing returns everywhere. You can also see this in the OpenAI post, where they write 7 times as much code to run 1.6 times as many experiments, and those additional experiments probably only result in minor improvements to model quality.
“Recursive” is a reasonable term because the generation N AIs will train the Generation N+1 AIs. The term “iterative” doesn’t reflect this nuance as well IMO.
Recursion reduces each step toward a base case: each step is defined in terms of previous/simpler steps, not more advanced ones. The "recursive" in "recursive self improvement" has things precisely backward. Iteration correctly describes a process where each step is the starting point of its successive step, so it should be "iterative self improvement" but I guess that didn't sound as cool.
Yes the exponential self improvement folks have never heard of an eigenvalue I guess. You can loop forever using output as input but at some point the result will stop changing (depending on the function)
I think the limit of what can be achieved with RL and synthetic data generation is better simply described as a leveling off of gains as you extract all the intelligence and knowledge from the original human training data.
Of course things will change at some point in the future as we go beyond LLMs, to build creative intelligence not just imitative/predictive intelligence, but right now these companies are stuck in this loop of building synthetic data and RLVR training from that, which means they are essentially building the "generative closure" of the original human training data - trying to squeeze all the juice out of it.
To go beyond this they need to add creativity of some sort to generate data that is not ultimately based on the original human training data. They could try something like brute force search (cf agent swarms/graphs), but this is just a more thorough way of exploring the search space defined by the training data - it may find you the "move 37" or low-hanging mathematical proof, but as Demis Hassabis has said, the goal of AGI is not to find move 37 but rather to create something capable of inventing as compelling a game as Go in the first place.
The "You can loop forever using output as input but at some point the result will stop changing" idea seems more related to dynamics, feedback and attractor states.
If you have a feedback loop of "recursively improving" by generating synthetic data, to train on, to generate more synthetic data, etc, then this is indeed a matter of looping "using output as input", and without any other system inputs it would indeed "converge" to some attractor states.
Eigenvectors represent fixed directions, not fixed magnitudes. From Wikipedia:
> More precisely, an eigenvector v of a linear transformation T is scaled by a constant factor lambda when the linear transformation is applied to it: Tv = lambda v .
In other words, repeated multiplication of an eigenvector by a matrix can still create exponential growth.
>AI "recursively" improving itself in some exponential fashion until there is a bright flash of white light
Sounds like repetitive stress to me.
>loop forever using output as input but at some point the result will stop changing
Running in place will eventually wear you out too. Plus with some things it can be difficult to know for sure if that's where you are at the time.
Even worse may be if you were almost running in place, it could be orders of magnitude more difficult to discern, especially if the scale was massive to an unprecedented degree.
For example, TSMC uses behavioral cloning to scale up human-bottlenecked parts of the manufacturing process to meet the growing demand, while automated research laboratories do thousands experiments in parallel to find better manufacturing processes.
The production bottleneck in a fab isn't the human workers - the process is mostly automated. The bottleneck more derives from how many wafers per hour you can process, which comes down to the etching process and EUV throughput.
ASMLs EUV machines are literally the most complex machine that mankind has ever built, which is why no other country, including the US, has yet been able to duplicate it. It's not just the machine itself, but a global supply chain of irreplaceable components such as focusing mirrors made by Zeiss to an incomprehensible level of accuracy - differences in surface height no more than the size of a hydrogen atom (or if you scaled the mirror up to the size of the country of Germany, then surface differences in height of 0.1mm).
Robots are useful to automate things, but they are zero help when trying to build tech like this that you are incapable of building in the first place.
The US has fallen way behind in manufacturing expertise, and no swarm of robots is going to help.
You've missed a part where it's TSMC that does behavioral cloning (to build more EUV machines). The full vertical integration is a bit farther down the line.
Etching a model's weights on silicon is another way to utilize non-top-notch tech-processes, while maintaining or improving performance. (and it suits robotics well)
TSMC doesn't know how to build EUV machines - they are stuck buying them from ASML like everyone else.
Putting a model's weights in read-only memory close to the processor is certainly a way to increase token/sec generation speed, but of course does nothing to increase intelligence. Robots aren't going to help though - semiconductor manufacturing is semiconductor manufacturing regardless of whether you are etching GPUs or memory onto your wafers.
Ah, sorry, it's ASML expertise that needs to be cloned and scaled up. I don't see how it changes things, though.
Robots don't need that much intelligence. High-speed joint control, "hand-eye coordination", the higher level tasks can be delegated to external models. Distillation already works quite well for isolating the required functionality.
I'm pretty sure ASML, and their supply chain, do have plans to increase production, as do the chip fabs - they all see the demand, while at the same time being leary of boom and bust which is the historical reality of the chip business.
But, the production expansion rate of none of these companies is being limited by lack of trained personnel, and if it were it would surely be faster to hire/train more humans since robots are still very far from human dexterity, not to mention intelligence.
Robots and AI are tools of automation, a way to replace humans with machines, but not all the problems in the world are bottle-necked by lack of humans, or the cost of humans.
Investing in training a person gets you one trained person. Investing in training an ML system gets you a cloneable ML system that can be scaled on demand much faster. ROI might change quickly.
Sure, but we're simply not at the point, maybe never will be, where lack of employees is the bottleneck to chip production. A fab takes billions of dollars and multiple years to construct - there are many constraints.
> What will prevent LLMs from designing robot control circuitry and participating in increase of chip production/design and physical experimentation?
Money, regulations, EUV machine lead-times, global helium supply, reality ...
It's funny that we've got the Dwarkesh contingent saying that GPUs will become infinitely expensive, and now another contingent saying that they will become infinitely abundant.
Even if compute were free, and/or the AI was so smart that it picked the right experiments to run every time ("make no mistakes"), you still have to actually train the model, which takes months, and if model Ver. N+1 depends on model Ver. N, then it's iterative regardless of how much compute you have.
Who's saying that compute will become infinitely abundant? "Singularity" is just a way of saying that known models begin to give absurd predictions. Anyway, intelligence is a way of overcoming obstacles. 10 million tonnes of helium is a nice head start and retraining models from scratch is not guaranteed to last forever.
AFAIK the notion of a/the technological "singularity" is a point in time where technology is building upon itself (RSI!) so fast, at an ever increasing pace, that the speed of change effectively becomes infinite and incomprehensible to humans.
The word "singularity" is presumably coming from math or space, like a black hole singularity where matter becomes infinitely dense and the known laws of physics break down.
reply