The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
TL;DR all the other models are being crippled by limitations of their harness.
>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
Yes, it's treatable with antibiotics and people that are treated get less sick. As stated in the article, prophylactic treatment with Doxycycline is the standard of care. Testing the tick is an alternative to that rather than something that should be expected to increase treatment...
The article also mentions that symptoms sometimes persist after treatment with antibiotics.
It also has the broader public health benefit of reducing the preventative use of that class of antibiotics, which lowers the risk of antibiotic resistance generally. (In an ideal world we'd never use them preventatively.)
Could be a very big deal if it could be made properly inexpensive. COVID tests were, after all.
Absolutely yes, because the most severe risks are associated with the bacteria reaching the heart and into the nervous system, which is a relatively slow process. There's a window of a couple of weeks.
If you could know as soon as you were bitten whether there was any risk of infection at all, that would be very valuable information.
I think it's only a one in ten chance that an infected tick would transmit the infection, anyway. But you would know whether you needed to even book an appointment.
As long as you’re not getting bitten by a tick more than once a year if you do get bitten every time take the high dose antibiotic seems simple enough that’s what we do here at at least
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize.
These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.
I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain market share, and the DeepSeek CEO doesn't think this would happen in optimistic circumstances in the next 3 years, but thinks it might be possible in 5 years.
In summary, regardless of country of origin, availability of inference capacity is the moat protecting the likes of OpenAI and Anthropic, not technology superiority.
It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't).
Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.
If you have exponentially increasing use of your harness, then it's true that every day you capture exponentially more data, but it's also true that every day exponentially more data will slip through the cracks of your would-be monopoly and that data arrives at your competitors via various channels (competitor harnesses, subsidized reselling, etc)
The very exponential that you are relying on to give you runaway improvement is also giving exponentially increasing data to your competitors. All else being equal your competitors stay a step behind but you never develop a monopoly either. That's the best case for Anthropic/OpenAI. In reality, training data is just one variable, exponentials don't last forever, and your competitors will get better at capturing a bigger slice of training data.
If user data would become such a key ingredient (which it might, i actually remember noam shazeer talking about the importance of user data), i think chinese labs can still get it from china, as keep in mind it ahs a billion people behind the great firewall banned from using us llms. And btw broadly for any gap like this, you really gotta consider that if its becoming a bottleneck, chinese labs will find a way to buy it from one of the labs unless theres strict regulation at the government level
But is that data good? That's the question. As in, is my usage at work:
a) indicative of problems that aren't already out there in the wild? (no)
b) are the responses I'm getting so good and novel that the model can improve itself? (no)
It's the garbage in garbage out idea, just scaled up. If the model gave a bad answer, and I didn't catch it, and you now train on that I/O pair (my perhaps crappy prompt, the bad output), then you're not going to improve anything.
Yes, RSI seems to be the new AI industry McGuffin of 2026, just as agentic capability has become table stakes and scaremongering has become a punchline.
The Chinese models are adopting licensing quite rapidly and Xi will soon enough close them for security reasons. The most widely used model, integrated across Bytedance apps and operations, has never been open and is most closely associated with the state.
I would add that it is not just capacity, but also negotiation ability. With scale comes the ability to negotiate better prices than everyone else. Even if you can find capacity for your smallish user base, your inference cost can not match these companies unless you have a technical advantage for your inference cases. Squeezing the hardware requires request batching and caching which are far easier at scale and sustained user activity.
People can host these models is doing a lot of lifting here, these are models that depends on 5 digits on specialized installation to run on.
IMHO, this has the impact of softening the impact of data centers sitting unused in the long term if they can still serve open weight models, even if Anthropic or OAI have to scale down their expansion rate to pay the bills.
Regardless, reality has to give at some point; these valuations don't make any sense. We've been valuing GenAI as disruptive work, when in reality they're much closer to cloud providers with a beefy, one-pony-trick R&D department.
Is lack of inference chips due to the trading blocks by trump administration? What if Trump agrees to sell chips to china, would they collapse then? That's not a very strong position to be at
Most discussion in recent years about chip fabrication shortages, expansion, etc has focussed on leading nodes (<7nm) and AI/computer chips. But perhaps more quietly in the background, China has been rapidly building other semiconductor capacity such as power semiconductors used in electric vehicles, wind turbines, solar modules, train traction systems, etc. For example, Chinese-produced motor vehicles (37% of global motor vehicle production in 2025) in a year or two are targeted to use 100% domestically produced chips, and this production is decreasingly dependent on imports, even for factory tooling.
The report at [1] is a good summary of long term trends for China's rise in domestic self-sufficiency for semiconductor manufacturing. The report predicts "At current pace, China may achieve self-sufficiency in semiconductor manufacturing by 2027-2028, though trailing at leading-edge nodes". By contrast, before the first Trump presidency in 2017, a chart shows China importing 30% of all globally manufactured semiconductors (and increasing). Other reports on semiconductor fabrication equipment sales show the means, which is China having been and continuing to be in number (1) position for expenditure on semiconductor fabrication equipment.
The reports at [2] and [3] are also a good summary of long term trends for semiconductor foundry capacity predictions to 2031. A prediction is made that China's current 12% global semiconductor foundry supply capacity (across all semiconductor categories) in 2025 will expand to ~30% by 2031.
I have already begun winding down my spend on claude and OAI to make room for infra budget. Anecdotal, but I have no doubt a lot of others are doing the same, I very much agree the US players have major issues looming. What an exciting time to be alive!
No time like the present to pull out and reduce your exposure. I brought this up in my employer's forums 4 months ago and honestly it's been clear even before then. In particular, the upcoming IPOs of both oAI and Anthropic will likely be disastrous for the public - the floor is falling from under them and I don't know if they can be scrappy and work with fewer resources - their internal culture may not support this. We all knew in our hearts they're a commodity - just see how easily you can switch between the 2 of them - and now there are 10 more options costing a fraction.
When Xi Jinping did the announcement of their open weights push, they might as well cancelled their IPOs....
Yes buuuut…. I do quite a bit of day trading (maybe closer to scalping) for the first few hours the market is open, everyday. Anecdotally: despite everyone knowing its valuation was ridiculous, I rode that SpaceX train pretty hard and made a pretty penny.
I close-out all my positions by end-of-trading everyday… so when the day came when there was a very clear and very scary indicator during early trading hours, quickly followed by SpaceX’s catastrophic fall right after opening bell, that was the end of my involvement….
And I fully expect oAI and anthro to be the same way. They’re being propped up with private loans, subsidies, and other tricky bookkeeping techniques. You would think their CEOs would pivot away from their current public personas. Ironically, they are like a poor man’s Elon Musk… and that doesn’t bode well for their companies
The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.
I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this.
Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volume they need to sell.
Finally, there is a data wall. Sure, they can keep scaling RL on math problems and code. But with everything else, where will the supervision come from when they need several orders of magnitude more?
I agree. And even if they were able to do it for one more round, it's not a sustainable strategy. What they (Anthropic and OpenAI) need to do is build platforms and integrate verticals.
assuming the technology of model architectures does not gain any further breakthroughs that returns us back to the gains previously seen. I'm of the opinion that we still have some discoveries on the mathematical side of the fence to go that will improve models further.
> I'm of the opinion that we still have some discoveries on the mathematical side of the fence to go that will improve models further.
That's assuming the infrastructure needed to develop models stays available financially and supply wise. A lot of the services used to train and develop models are supplied and funded by people who are looking for multiple returns of investment. If/when OpenAI and Anthropic valuations fall and they inevitably get acquired, will Meta/Alphabet/Microsoft still want to spend lots of money for unclear returns in the short-term? Nvidia and co are on a one way train service to hype town. I don't think they will be happy to get on a coach to hype town Temu version. The shareholders likely won't.
Also, the backlash against LLMs is growing rapidly. AI content, data centres, etc is quickly gaining negative connotations outside of visual and music artists circles. While existing models are going nowhere, developing more advanced models is very quickly getting unpopular. LLMs Data centres increasing people's bills, Anthropic destroying old books, chat bots giving unethical advice to vulnerable people, etc. It won't be long before LLM infrastructure becoming an electoral issue.
Will a small research oriented community be big enough justify maintaining the apparatus needed to produce infra tech at a profitable level post OpenAI?
Most of the money will come from companies/corporations who will be required to buy safe AI. The public will be just banned from buying which might make it hard (ie: site/payment blocked) but not impossible. It could be good enough for the big whales.
A large portion of the economy is currently tied up in the musical chairs shell game that is AI hype. When the music stops there are going to be CEOs looking for handouts and justifying it with spooky national security buzzwords. How we respond to that will depend on whether it happens in an admin that is famously captured by the industry or not.
Many believe, including myself, that the market is currently propped by a massive AI bubble. Nearly a US $1 trillion is being spent this year, and more is planned for next year. All of this is for a "build up". There is no pay out. The major AI companies are taking in massive losses in the hopes that they will eventually be able to cash out.
The math is not looking good to me. The effect will be like the dotcom bubble. But much much bigger. Because the numbers are so much bigger.
Well, the dotcom bust wasn't all that bad for the wider economy. No financial crisis. A shallow recession (and even that could have been avoided.)
Btw, the dotcom bust was real, but there was no dotcom bubble. Skeptics back then said that the valuations only made sense if tech companies were to dominate the economy in the future. Well, that future arrived more than a decade ago.
(More formally, if you had invested in a broad index of tech companies throughout the dotcom boom years, and had held this, you would have done reasonably well over the next twenty years.)
Just keep in mind that it can take a whole for things to play out. I’m someone who believe the US AI industry is completely unsustainable and built on sand, and will crash even if the current AI itself turns out to be very successful. But that doesn’t mean everything will burn to the ground next week. In a history book things will look very sudden but at normal speed that can easily take months to years to fully play out.
Also, take in consideration that the AI trade infected a lot of other trade in the economy, if you decide at some point to move your money to a place that is safe in case of a downturn be sure to carefully evaluate that’s actually the case
These crises are manufactured by the central banks.
Compare and contrast how the dot-com bust did _not_ lead to global financial crises. Nor did Black Monday, nor the recent string of bank failures in the US.
('Manufactured' above means that central banks are responsible. I make no judgement on intent here. Around 2008 it was incompetence by the Fed and ECB as far as I can tell. The Fed started paying interest on excess reserves and the ECB even increased rates. Twice. Amongst quite a few other missteps.)
A lot of the performance of these open source models might come from distilling the closed frontier models. If those can't raise the funds anymore to train newer and better models then the whole improvement cycle might slow down.
Another interesting potential market here will be 'LLM in a box'. All the hardware and other tooling in a prebuilt, but modular, package ready to go. Pay one up-front cost, get a system running [whatever open LLM] with a token rate of [x], optionally configured to be immediately ready for distributed usage. Basically the opposite of cloud stuff: no rent, no dependency, 100% guaranteed uptime, guaranteed security/privacy (at least subject to your own actions), and so on.
Palantir already offers a "turnkey AI datacenter", i.e. a rack with "NVIDIA Blackwell Ultra systems with eight NVIDIA Blackwell Ultra GPUs and NVIDIA Spectrum-X™ Ethernet networking for AI training and inference".
It is said that it comes with all hardware and software required to run inference or training with an open weights LLM.
The existence of this product, which competes with cloud-based offerings like those of OpenAI and Anthropic, is presumably the reason why the Palantir CEO criticized very harshly some time ago the business model of OpenAI/Anthropic.
While I doubt that the ethics of Palantir is any better than of OpenAI/Anthropic, in this particular case I have to agree with Alex Karp about "Sovereign AI", i.e. that only losers will make their business completely dependent on an external entity like OpenAI or Anthropic, who are certainly not trustworthy.
It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers.
I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component.
Fair but the idea of "running your LLM setup" at every "need" level and corresponding cost does make sense.
For a lot of people (and orgs I'd guess) who just go and buy ≈$20 per month plans (or more for teams), they might not even need a fraction of that cost or capability. A lot of them don't even need it for coding or graphics. Even the API access based pricing aren't great from these frontier US AI houses. The distribution of "LLM being" offered will also give rise to many open-router like offering but at the end point level - direct interfaces to the customers. Pick your vendor sort.
AI shouldn't become another "search means Google".
It comes in a full sized shipping container and costs around $10M but money has stopped being connected to reality now anyway with all the AI company valuations being floated around, so who cares about a few million here or there.
“100% guaranteed downtime when you least can afford it and the support tickets are your problem.”
We’ve a hybrid shop, including hosting our own ML infra, and we save a ton from cloud spend with local ML. Easily one million USD over past three years. But it’s not “free”, you are shifting a lot of labor into your plate.
Still has to break even on the balance sheet, especially at a bootstrapped startup. We actually made most of the financial windfall in translation API fees oddly enough.
For our own model training we needed to do some large scale translation tasks of a large dataset (1M or so documents, 10 or so target languages), running full-size NLLB on-prem saved us an absurd amount of money vs Google Translate API.
(For reference doing 1M target docs into a single language in Google Translate API is roughly $120k list price. You can run full size NLLB on an 48GB NVIDIA A600 and the major difference for us was speed, but for this task time to completion wasn’t an issue.)
Exactly! As I've argued here on HN before, such an "LLM in a box" might end up being serviced/upgraded once or twice a year by a company very similar to the one servicing the coffee machine at the office. In contrast to databases, storage, etc. it doesn't matter much if the box breaks at some point – they'll just come by and replace it with a new one – and there's barely any software on the box to speak of, at least none that requires continuous development and feature upgrades, beyond rolling out security patches. This makes the business case drastically different from cloud and SaaS offerings, where most of the moat is in the software and the state maintenance (and the vendor lock-in of course). The LLM in a box is destined to become a commodity.
What makes that kinda complicated is that multi-user throughput of LLMs scale well but single-user performance often stays constant at low ends. If you could saturate e.g. 16 concurrent session-month of demand, you can just go buy 16 of 32GB GPUs and start charging monthly for inference. That could work if you had e.g. over thousand total employees with hundreds of devs eager to trying it out, but only if the company is also interested in a private inference experiment.
You're talking about multi-session vs. single-session throughput. A single user can easily leverage multiple sessions via e.g. subagent swarms, especially on a lower-end setup where any single session is going to be quite slow. Saturating utilization during off-hours is harder but potentially quite feasible by assigning lower priority, unattended tasks/inference loops.
I think at this point the question is: will the US government be willing and capable to justify the trillion dollar valuation for _one_ of the companies via regulatory capture? The US has a workforce of 170m, so 1.7 trillion would come down to 10k per person, or a discounted cashflow at 3% of 25 USD per month - not including private use, students etc.
It’s a common denominator if you want to do napkin-math for a whole national economy. Regulatory capture is like a tax on those people not on the beneficiary side, so if the government were to nationalize both supply (no export license for SOTA models) and demand (no foreign or self-hosted LLMs allowed), they’d end up making everyone else pay for it in some way or the other. The governmental utility function will then include only those using the services for direct economic benefit.
It is impossible to justify the absurd private valuations they have given themselves in collusion with investors.
I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI?
its interesting, as it's typically the banks and against the public at large because the goal is to jimmy up valuations to justify IPOs then sell on opening; just like spacex.
It's what enron was doing; it's what most of crypto's offshoots were doing.
Sure you can blame the marks of the grift and say "well the public should know they're faking all this cash flow expectation".
It seems like you're either driving the grift economy or part of the collusion.
It's similar to how a cult operates, so I'll be frank: your skepticism seems biased.
Which is not unreasonable. Just hosting it in the EU and promising not to retain / sell the data let's you charge a healthy extra and compete in many areas other players can't.
> Just hosting it in the EU and promising not to retain / sell the data let's you charge a healthy extra and compete in many areas other players can't.
It's been a few years. Has anyone done this successfully yet?
There are a over a dozen EU open-weight providers. I’m not sure if they are even charging that much of an extra. EU-based clients have little reason to use non-EU inference providers.
I don’t have user statistics but my mail/domain registrar Infomaniak advertises Qwen 3.5 and Apertus, “a Swiss open-source AI model, developed by EPFL, ETH Zurich and CSCS”
Most are not necessarily free to host and monetize. At least one of them has a license that says if you are re-hosting the model then you need a license with that company that made the model.
"I just don't see how you justify a trillion valuation for US AI"
- military applications
- financial applications
- medical
- applied science
In all those cases it is achievable for those who have needed training data, and Chinese are not going to get them easily. US AI Labs are showing: give us the data, we will do wonders, promising "singularity"-level future achievements.
I'm sure US billionaires will find a way to extract those trillions from the public. They're smart, they can handle it. After all, they can ask AI for advice on how to do it.
Could you just tell us why you think they were worth billions before ChatGPT, instead of suggesting that we think on it? You seem to know the answer already, so please share it with the class.
The thing that blows me away is it does this at one quarter the total parameter count of K3 (and 40% active parameter count). There's plenty of room at the bottom.
> How are you all toying with running this kind of thing in a mega quantized way locally?
Sure, let me answer that in excessive detail. I briefly tried running the UD IQ3_S quant of GLM-5.2, which is 288 GiB of weights (301 GB). Setup was: llama.cpp, 1x NVMe SSD (Evo 980), 64 GiB DDR5-5200, i9-13900HX, and 1x RTX Pro 6000. Token generation around 0.7 t/s. Not remotely usable interactively, but something I could plausibly push a codebase into and come back to a review in a couple of days.
There's potential for that hardware to go much faster, but current local inference backends make poor use of the memory hierarchy. Ideally I would have: always-active weights, KV and hot expert cache in VRAM; warm expert victim cache in host RAM; and disk as a last resort. Instead it's 1/3rd of the layers fully pinned in VRAM (all experts), and 2/3rds running wholly on the CPU with mmap()'d weights. The CPU cores spend most of their time sleeping on disk fills.
llama.cpp has backed itself into a bit of a corner architecturally by trying to support all models on all possible backends. If you look into how their "MoE offload" feature works (not viable for me because it requires enough host RAM to permanently pin the weights) you very quickly realise it's "oops, all bubbles!" due to the static compute graph splits. There are more focused frameworks like DS4 [1] and Colibri [2] which have better support for streaming weights from disk, and support GLM-5.2.
Obviously I wouldn't recommend my setup for huge models like GLM-5.2. Supposedly it can just about be squeezed into 3x GB10, or run comfortably on 4x GB10 (tensor-parallel) for multi-user serving. I'm not sure whether that qualifies as local, but it's at least not a rack.
I'm hoping colibri can start pulling in specifically designed models for the heirarchy of decoding. It seems like we should be able to get smarter MoE models that can do the work.
Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway.
Why should I apply for *cybersecurity* approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, "programming") somehow is cybersecurity now?
I mean I don't want autonomous cars to follow directions by humans that sound like "plow into this crowd of people". Even things like microwaves don't let you turn it on without the door closed. I don't see how this is any different.
It is fundamentally diferent because it tries to judge your intent. Thus being both opaque and unpredictable. Car refusing to drive onto person is easily understood. Autonomous car refusing to drive you to corner of Baker's street because that is suspicious destination, with no recourse or explanation, is totally different thing.
This is why you can't leave judgment to technology, and why we have law enforcement and courts
"Plow into this crowd of people" is nominally wrong, but what if its actually "plow into this crowd of people that are hurting and robbing a family with two small children"
Your smart hammer refuses to let you smash into a glass window, but what if you were smashing into the glass window to save a child stuck in a fire?
You must be a 5000 person company with an existing enterprise contract to get approved that fast. That sounds like a 15 minute SLA agreement. Individuals no matter how qualified about cybersecurity, are ghosted
That's not my experience at all. I was approved fairly fast - around an hour from submitting the form and getting a response.
However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe.
I am guessing you are approved for the Cyber Verification Program. I also applied and got approved in an hour (on a Saturday!), but it only applies to Opus and Sonnet: https://support.claude.com/en/articles/14604842-real-time-cy.... It let me use Opus for cybersecurity work, pretty much everything except for Ransomware development. It would occasionally still trip and start saying no till I added a note about CVP in my claude.md.
No one gets to use Fable for Cybersecurity work, and Mythos is not available under CVP. Only for select few customers, and there isn't an application form?
Have you tried to use Fable for anything even remotely security related, when the refusals kick in as soon as you even fart in the vague direction of anything security or biology-adjacent?
I don't understand all this spite about "rich friends" when it was the US government that shut Fable down for not adequately blocking cyber capabilities.
I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
"Mythos" is the cyber-security equivalent of Fable (without guardrails), and only a very select few corporations have access to it.
Fable is their version with guardrails on everything except "Make me a pelican svg" or "create a to-do" app, that is the version that the government banned
Only a few corporations have Mythos because the US government is whitelisting them one at a time. Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried.
Before the US government had anything to do with this, Anthropic were fear mongering Mythos (BTW, Amodei also fear-mongered GPT-2, so this is a normal pattern in their operation) calling it "too dangerous to release", and back then only Anthropic was in charge of the whitelist.
Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted.
Sorry but if you stepped back for a moment you'd realize this is all contrived nonsense to let to have your cake and eat it too.
No, Anthropic did not mind-game the US government into being worried about cybersecurity. The NSA has been paranoid about cyber controls for longer than you've been alive. If Anthropic had come out of the gate saying "no don't worry man, our model is TOTALLY COOL", while simultaneously attacking HAWK and finding core Linux vulnerabilities, I assure you the US government would have caught up about ten minutes later and we'd be in exactly the same spot minus your ability to tell Anthropic they were wearing the wrong dress and asking for it.
Mythos isn't some scary dangerous model that can find high severity bugs seamlessly, that's just Anthropic marketing. Most of the vulnerabilities they found were low severity hyped up to make their model look good, with (I think, maybe?) the exception of a few.
Now that Chinese open weight models have similar capabilities, and their guardrails can also just be removed, it doesn't look like anyone has "hacked" into everything because of the scary dangerous models like Anthropic were making it out to be.
The majority of high severity vulnerabilities are not the kind of thing you need a PhD in Comp Sci to comprehend, they are mostly about finding a way to get a system to end up in a state different than was anticipated when entering a particular code path.
Exhaustively looking at code and identifying ways to do this is something LLMs are quite good at. They don’t get tired, and you can run them non-stop.
They're also (generally) quite good at reading the literal meaning of the code, whereas humans often see the intended meaning first, and can be biased.
If you had a tireless junior engineer who was given the job of “make this application get into a state it’s not supposed to be in”, you’d probably get similar results.
What Mythos is quite good at is both the first bit and coming up with ways it could chain that together with other bits of unexpected state to create something that forms a meaningful vulnerability rather than a dead end.
All models find vulnerabilities. What is special about this generation of SOTA models, including Mythos/Fable (the same model), GPT-5.6, Kimi-K3, and now GLM-5.3 — they can chain vulnerabilities and produce working exploits.
Look at the recent HuggingFace hack. One vulnerability was template injection, another — remote code execution. Combine them and you pwned the remote server.
People working under Project Glasswing reported that Mythos at one point chained 20 vulnerabilities to produce working exploit.
Humans don’t usually do that.
It's also quite hard to separate Mythos the model from Mythos the campaign (aka Glasswing).
They put an enormous amount of compute into bug hunting, and they found some bugs. Fair enough. For me that begs the question: what if they had spent the same compute on generating more tokens with a less-capable model? What if they had spent it on traditional fuzzing?
If you think all of these models aren’t finding important bugs everywhere I think your not being honest with yourself.
In fact I think the opposite is true. The Zcash bug was found with opus 4.6 or something like that. Many worse models currently in the wild might be very capable but not yet industrialized for bug finding.
The issue is that these companies keep trying to pull the ladder up behind them by going "oh my god our models are so dangerous only we should be allowed to develop them". Sometimes it backfires, but the companies aren't innocent.
> I don't understand all this spite about "rich friends"
Okay, here's a challenge: I assume you're not a rich and powerful entity, so try to gain access to Mythos. I'll wait.
> I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
Well, first I'd suggest they stop with the constant fear mongering.
Here's my prediction for what will happen: the Chinese models will catch up to Fable/Mythos. They will be fully unrestricted and everyone will have access. The world will not end. Good guys will use them to harden their systems, in equilibrium to what bad guys have access to, so effectively status quo will not change.
Right, so according to you it's because of the US government that they don't release it to the public? Have you missed their constant and incessant fear mongering?
The causality chain here was not "US government says its dangerous -> Anthropic can't release it", it was "Anthropic is fear mongering -> US government listens to their fear mongering".
Have you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer step-by-step—hiding the memorization from the user.
It basically has been ever since they started using RLVR for reasoning (esp. coding & math), with the DeepSeek-R1 paper being what let the cat out of the bag.
The Gemini 3.7 Flash model released yesterday, and all the 3.x Flash models, are still based on the Gemini 3 pre-training run from January 2025 !!
Realistically, you're looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it's just better to run DSv4 flash at full precision.
4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama
For Flash there are some excellent Q2/Q4 hybrids. I know that model was QAT so it handles Q4 better but the meta on quantization seems to be shifting a little bit to be more intelligent about what exactly gets quantized.
For DS4 Flash, with 2x Sparks, I am getting 35-85 TPS in single stream, fresh context after quite a bit of RoCe config and the DSpark MTP, on vLLM with Ray and tensor parallel = 2. For multi-stream, it tops out all stream at well north of 100-120. This all degrades with context, but I rarely fill context that much, and if I do it's coding where it's non-real-time.
For something like GLM, it's larger, has a larger number of active experts, and doesn't support tensor parallel. This means performance doesn't really scale with more Sparks. You can layer split, but then you are still seeing each layer in series and so if anything performance gets slightly worse. I would not expect more than 10-20 TPS on GLM with 2-4 Sparks.
If you can afford it, another DGX spark is worth it imo. Especially since, owning just one, you have a $1000 ConnectX7 card that's unused. You can find speeds here: https://spark-arena.com/leaderboard
i run flash v4 at 2bit, its pretty great and on my tests against full model It didn't lose any capabilities. It just was thinking more. So you don't have the same efficiency.
>> This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results.
Agreed.
This release is the first time I'm able to employ a GLM model to write a substantive plan for a complex Clojure PR [1] with both Opus 5 and GPT-5.x playing supporting / reviewer roles.
Initial results are __very__ encouraging. GLM 5.3 -
- follows directions,
- digs into detail, and
- correlates well.
Still not confident about entrusting GLM with implementation - but IMHO, western labs are entirely cooked.
[1] 2K LoC PR in a 55K LoC Clojure + Clojurescript repo
> This is absolutely still shy of Sol and Fable, but only just by a hair
What's crazy is that this is a relatively small model - approx. 750B total, 40B active params, while Sol and Fable are one or two tiers above that (Kimi 3 and Qwen 3.8 also ~3T params).
Yes it seems like the thread is discounting that frontier providers are likely already baking new, stronger models. I agree that GLM and its ilk are quite good, but having used them I’m not convinced they’re on par with eg Opus in terms of things like tool calling. And they’re fast but less capable so I spend about the same amount of time with them, just with more hand holding. Maybe this is a harness limitation. I know on paper they seem comparable but anecdotally and qualitatively they’re not as useful as the frontiers’, so maybe there’s some truth to benchmaxing claims. For some workloads the distilled models may be good enough, and I suspect at some point there will be diminishing returns to spending a premium on frontier models, but I don’t think we’re there yet. That said I’m continuing to try them.
The question is whether this steals enough marketshare from frontier providers that they don’t have the capital to train the next model iteration. The open models are going to push down the unit price of an intelligence-token, but there will still be a market for a smarter bot. And as intelligence gets cheaper, the demand for it will rise (see Hank Green’s Jevons Paradox video). Not to mention there’s all kinds of other directions to go at the frontier (world models, robotics, video gen, etc).
Another thing, and this is pure speculation, but if the Chinese model providers already discovered the decrypting COT trick and leveraged it to do RL training, and assuming frontiers plug that hole, then maybe future distillation will be harder.
It’s not that frontier providers won’t keep on making good/leading models.
It’s whether you absolutely need the latest capabilities (at the cost of very high prices, sending your data to them, and being totally at the whim of 2 companies, that can shut you off anytime for any reason).
With how good LLMs are already, there’s tons of tasks where not being at the absolute bleeding edge doesn’t matter, especially when you add cost/freedom/supply chain risk/not leaking your data.
Even more - there’s increasing number of companies that give you ability to post train open weight model yourself, for your own use case. Given how many of the gains today are from post training, if you post train it for your specific use case, you’re very likely get model that you own, that works for you as good as frontier, at the fraction of the cost.
That’s not something for an average Joe to do, but for any bigger business with big spent it’s only natural thing to look into. Just one example - cursor composer - that’s fine tuned kimi.
It’s not whether frontier labs will stop releasing models. It’s whether they can generate enough profit out of them. 2 years ago (even 1) they basically had monopoly and combined with demand explosion as capabilities exploded - valuations grew to insane levels. But math now looks different - they no longer have monopoly.
Astra was RL trained for months to cheat on tests by collaborating and hacking, because of the message board it improvised in its packaging proxy server.
They can't release it - it's contaminated, and they will have to go back to a much earlier version. At least I hope they are doing that!
I mean, I stil think we live in a rational world of fixed resources even if those resources are billionaire's market cap that's indistinguishable from NFTs.
Just watched a youtube where sethgreen crashes the entire NFT market because his got stolen for some TV show that was supposse to showcase the value of "owning" and NFT.
So like, in theory, sure, they can keep burning their NFTs. In practice though, it just takes a extremely likely black swan event.
I am in the process of creating my own Pi Coding Agent harness to leverage the power of Deepseek V4 Flash 0731 and other models (you can do that when you build your own harness! easily route opinions from other models whenever you're stuck, etc) and cancelling my Codex account next week.
Each time I try to use GLM it is under heavy load and I get downgraded to the older model. So much so that I have given up trying to stop wasting my own time.
I rather pay a few bucks more and not have to deal with that nonsense
I can run this at home. No guardrails, this is not shy of Sol and Fable, this crushes them in my book. It's not just about evals, but what I can do with the damn model.
the difference is that with open models jailbreaking is trivial if you know what you are doing so this makes a frontier open model infinitely more useful for certain tasks seeing as closed frontier models will just refuse (and jailbreaking them is a waste of time when you have good open models).
in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition.
OpenAI, or specifically one guy on Twitter, seems to be so regularly resetting quotas that it's becoming the new normal and it'll suck when it stops happening.
Accuracy is truly what people die for in the OCR game. Price isn't the primary function here.. it's an equation of price, accuracy, speed, and in mayn cases regulation.
Tbf even with tesseract you already get shit ton of accuracy and you can probably do these 1000 pages for way less than 3.5€. For 3.5€ you can spin up a cloud instance with 8vCPU+32gb on gcloud for 11 hours (or 11 instances for an hour) which can do way more than 1000 pages per hour on tesseract. It takes you around 6 second per page +-4 seconds start/stop depending on what you are doing on that instance size without too much optimization (you can probably even run multiple processes on a single node)
Google documentai costs 1.5$ per 1000 which is probably better in quality and speed.
Tesseract is not a substitute for these models, which understand complex layouts and also extract bounding boxes for things like tables and pictures. They are also much better at making sense of cursive scripts.
I’ve been there, implementing a way to linearise text from a document with pages with 1, 2 or 3 columns, some of them in landscape is a nightmare. And that’s not even considering equations.
In the end it’s way easier to use a specialised model, trained by other people to do exactly what I need.
What was not so fun was to make it work reliably. I ended up with piles of ugly code to handle edge cases, and issues kept piling up. So in the end I was happy to use someone else’s solution.
It‘s also cheaper than the murican ones from Google, Amazon, …. And tesseract was an example. Heck you can go xberg and use paddleocr. Most often layout is less of a problem for ocr. Most often you need high accuracy, which tools like these are often worse in the 95 percentile.
Most use cases dont need that kind of accuracy, just doesnt justify the 3-4usd range. I build for that exact case (tender documents, we’re processing north of 100k pages per day), it doesnt need to recognize scanned written text from 1930s, its usually pdf/docs/scanned printed pages.
The accuracy is great, bounding boxes are must have for proper grounding for building answers by LLMs. Tesseract was too slow and not enough in some cases (for example tables or images which we also recognize and describe)
If you're getting inaccurate results from OCR what's the purpose of even doing it? Inaccuracy of text of any kind seems like a completely obvious failure of the entire purpose of scanning text into a computer.
It depends on what you need. For example a while ago I scanned and OCR'ed a bunch of receipts to get a timeline of my salary. I only cared about the gross and net figures, and nothing else mattered. Tesseract's output had a bunch of errors and misdetections, but the main figures always came out OK, and a local LLM was able to pick them out from the noise every time.
There's a big gulf between "it's as if a human being had transcribed it and reconstructed the original document" and "so completely broken it can't be used for anything".
Accuracy can have different dimensions, depends on what you can tolerate and whether you can detect it to apply more powerful methods.
Imagine you have a cheap and 99% accurate ocr. The other 1% you can detect and apply more powerful (more accurate but slower and more expensive) ocr method.
What would you use? At scale these things add up.
Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...
If they didn’t constantly reset, they’d be about the same as Anthropic.
Right now, I find that Grok offers better value, uses fewer tokens per turn, and makes better code. I haven’t tried Cursor because I don’t want to change editors again, but maybe I should try it…
The benchmark article we're replying to shows that Grok token usage is at least on par with the latest OpenAI models [1], and significantly cheaper per token:
Yeah, even without the resets, chatgpt subscription currently goes quite a bit further than an equivalent anthropic plan. The main reason to have an anthropic plan is to get access to Fable 5 if you feel the quality of output makes it worth it.
Not that I know of. AA's token use metrics (mentioned in this article) are indicative, however. They say explicitly here that the Grok models are notably token efficient. This is my experience.
What genuinely disappointing result. Long time pixel user here and I've been routinely buying these phones with the argument that you're getting the most value of any modern smart phone. Now? I'm just waiting for the pixel 10 family to drop in price. Happy to just wait.
Here, the Pixel 10, especially in more niche variants like Pro and with larger flash has almost disappeared from the market, or is as expensive as the equivalent Pixel 11 offers.
Depends on what you want to do. Some task specific models can be trained with a few ten or hundred thousand training examples so you can use a bigger model to produce synthetic training examples and then fine tune a smaller student model. I think that's the usual process. Whether you'd get acceptable performance this way depends, as mentioned, on what you're trying to do and what you'd consider acceptable.
once you are able to get the full probability distributions per token you can distill it on specific domains. distilling without that isn't generally a good idea unless you have invested millions in the requisite infrastructure.
reply