There is going to be a flurry of this sort of stuff as the AIs get smart enough to find them. It will naturally die down as the legitimate ones are fixed. Yes, there will always be some level of this, but I’d expect it to be low and the exploits found to be increasingly complex. This is a time of transition.
> a flurry of this sort of stuff as the AIs get smart enough to find them.
I really think this characterization is misleading. It's not "getting smart", only more tailored toward a specific usage, better curated dataset, better harness, better prompts, better labeling of results, documentation of failures and success, etc.
The outcome is (hopefully) overall better but this anthropomorphized wording makes it sound like AI itself is somehow changing or evolving. No, both academia doing fundamental research, industry making it available commercially, and finally security researchers making the entire tooling and process packaged as a service are actively shaping it to make it better. There is no "it".
Long-term memory, for one. (reliving your entire life every time you do an action isn't memory). Creativity in new areas without training. Children at school are capable of "discovering" math solutions/methods that are known to others but hasn't been taught to them.
There's nothing intelligent about a math processor, even if it's automated.
Do you consider the protagonist of "Memento" to lack intelligence, then?
> Children at school are capable of "discovering" math solutions/methods that are known to others but hasn't been taught to them.
LLMs have already done that one: A chatbot’s result for the 80-year-old “unit distance” conjecture is the first AI proof that would likely be published in math’s top journal if humans had done it alone
My point was more about agency and anthropomorphization than the definition of intelligence, which is why I didn't just quote "smart" but rather "getting smart".
A future AI may be intelligent, but LLMs are clearly not. They have no agency, no ability to reason, and no world model. The most effective way to use them is to treat them as next token prediction machines, because that’s what they are.
edit: downvotes but no rebuttals. feel free to show me where the agency, reasoning from first principles, world model etc exists. or you can ask an llm and they'll tell you they don't have those.
I think you are giving the word "smart" a meaning or implication that it no longer has or is used with. It is common to say Google Search or Siri got smarter/better or dumber/worse, so I don't see saying LLMs getting smarter is any different.
> It will naturally die down as the legitimate ones are fixed.
Seems like we're already in the middle of this phase, but rather than dying down, the 'reports' have just gotten more noisy and obtuse, making it more difficult to establish the actual degree of threat / attack vector.
I mean. Makes sense until adversary states start walking through the same doors you’re using. At which point you might regret that maintainers are too flooded to deal with it.
Assuming, of course, said state agency is operating under sufficiently strategic governance and management…
I actually do not expect it to die down: as legitimate ones get fixed, LLM-based tools will continue finding other unfixed non-issues repeatedly, and overeager "researchers" will continue reporting them.
Perhaps, but as the AI analysis becomes part of the release process (or even the CI process as prices fall), you’d expect those new issues to be caught before release and fixed. We’re seeing them caught post-release for now because the code is older than the AIs, so we’re catching up.
Yeah, I am seeing this as well, especially as people use AI to code review stuff more so this sort of thing slips through. On one very large project I am looking at, it's already becoming harder to find issues.
These people whinging about slop don't realize everything that doesn't come from a credible source gets ignored.
Credible people are using AI and once these issues are fixed, it will die down.
The threat of AI zero days will persist though, but they will be much more expensive and subtle to find.