I'm kind of fascinated by how many of the same audiences who are highly skeptical of OpenAI and Anthropic are the same people running straight to other country's models.
The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?
This idea that the US labs are just directly committing fraud against effectively every major US organization is the most tinfoil hat thing I’ve heard in a long time. Training on excluded data would (eventually) be trivially provable. Forget loss of trust; this would make the labs a defendant to the most legally well-resourced organizations on the planet.
they knowingly and willfully broke copyright laws the world over in the faces of some of the most powerful corporations , but somehow some magical post-training method to discern provenance is going to be the legal gotcha that bothers the AI groups?
That's very defeatist. Do you have any concrete reason to think the major providers are lying to every one of their business/API customers about not training or storing the data? The business loss of trust would outweigh any benefits of the data.
(And if they freely lie about such things, I don't know why they would bother taking the PR hit when they announced fable had temporary data retention for their abuse prevention)
Because they know that in the end there will only be a few winners, and if you are one of them, you will settle even if its for billions, and if not it doesn't matter because you will be bankrupt anyway.
And besides various companies have been caught ripping torrents and other copyright data, what makes you so sure that same companies wont rip your data too.
Thats on top of 50 to a 100 years of companies straight up breaking the law to get ahead. Various ubers and food delivery app being the latest example.
I dont trust a lot of companies that i have to work with in some form or other anyway, with Oracle being on top of my personal shit list, followed by Salesforce and Broadcom.
And I dont think most of the AI companies are more ethical than any of the above.
AI labs are limited by the available training data, the best way to improve model performance is more and better training data.
The AI labs and the downstream companies that sell training data to them vacuum up everything they can.
Illegal residential proxies (botnets) that once have been used by hackers and scammers are now used to vacuum up the Internet.
They are now vacuuming up antique books that are practically useless.[1]
In face of this is is unthinkable to me that they are not training on API data.
> The business loss of trust would outweigh any benefits of the data.
The loss of trust is already here.
I know of one German company that uses AI only in areas where they have to compete with (foreign) startups. For their core business and everything else they are waiting for an on-prem solution. Apparently Microsoft can provide on-prem GPT-5.
> That's very defeatist. Do you have any concrete reason to think the major providers are lying to every one of their business/API customers about not training or storing the data? The business loss of trust would outweigh any benefits of the data.
At the same time if I was one of those large orgs and was running out of training data and falling behind the competitors, I'd probably have to stretch every definition under the sun, like what counts as "metadata". From a zero sum game perspective, it doesn't make that much sense for them NOT to train on your data if the consequences upon (non-guaranteed) discovery seem largely inconsequential when everyone just wants the best model regardless.
In this situation your comment is correct, but in many others the open weight models are hosted on different platforms like DigitalOcean etc that have different privacy policies.
When using (not this 'Stealth mode') open weight models you can choose a hosting provider you trust.
tdlr, chinese open models are more private and secure.
openai and anthropic plans steal your data. if you opt out they still log it, and send it to moderators to view if it gets flagged.
the other models are very often hosted by western providers with much stronger privacy and tighter contracts that the other subscriptions won't offer. they don't train, don't log and don't send to moderators.
The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?