Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

While we're dealing with the same issue at work, I sometimes still wonder exactly who these scrapers are. OpenAI, Google, Anthropic and others are normally fairly well behaved (Minus Anthropic attempting to hide behind a browser-for-hire company). Mostly you can get IP range and user-agents for the large players, while it problems mostly stem from bots pretending to be Chrome.

Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation. I also don't recall ever seeing Grok IP ranges or a specific Grok UA, but that doesn't mean that they're hiding, perhaps they're just not interested.



You need to account for the fact that a lot of scraping is delegated by the big players to smaller players who can take loss of reputation and can use questionable methods (residential IPs). Some of the scrapers from these AI companies have been written very poorly from performance perspective.

From what I understand the pressure has created even Google to be a lot more aggressive than what it was before. Not entirely sure but I believe google has two categories of scrapers, the regular one and a new one for AI


if they have agents write the scrapers you better believe they're playing every evil trick in the book


So it could just be 1 guy with scrapy and infinite vc money


Still not sure about the "who" question, but here's the scale of money changing hands:

> Residential proxies are everywhere, so why did proxy DDoS attacks mostly come from the U.S.? The answer is money. If you are committing fraud or circumventing content restrictions, a Russian IP address gets geo-blocked instantly. A fresh U.S. residential IP address (especially one behind carrier-grade NAT and harder to block individually) is “gold.” Customers pay up to $95 to lease a single U.S. residential IP address for 2 weeks (versus $0.30 for an Eastern European IP). Compare that to your own ARPU per subscriber and sit with it for a second. When an individual IP is worth more than the customer relationship behind it, you don’t have a technical problem. You have a market problem.[1]

1: https://www.nokia.com/blog/one-year-later-the-residential-pr...


Here is the fun part, the same companies invests millions of dollars to protect its infrastructure and services from scrapers/bots AND they also invest millions of dollars to circumvent the bots, captchas and rate limits.

Unfortunately you cant have it both ways which makes it hard problem to solve for everyone


if they pay 190/mo to proxy, i should just sell my ip direct -> profit i pay less than half that for fiber , should get a few more isp accounts, and rake it in.


One of the defining characteristics of AI turned out to be utter and profound facelessness. The scraping, the data centers, the slop content, it's like it all comes from thin air. Nobody is talking about who is really doing all this or why, because it's always just coming from... somewhere else outside of our place.


And then it's pushed on the end users with the same amount of explanation and reason. You must adopt this now. The FOMO is insane.


> The FOMO is insane.

do we have the same understanding of what fomo means?


I'm talking about managerial AI-induced psychosis and/or FOMO. Make the underlings use it so they can tell the shareholders that AI is making great strides.


You aren't stuck in internet traffic, you are internet traffic.


Well, some websites claimed China is behind it, which could make sense (I would not know either way). At the same time, though, I kind of doubt your carte blanche here for all those companies. Why would you think none of them are responsible for the AI slop spam?

> Our largest offenders seems to be mostly limited to South-East Asia, so probably mostly Chinese AI projects, but that's speculation

Ok. So you also don't know. Well, I don't know either, but I don't make a speculation by claiming x, y, and z companies to be exempt. In my book they are all responsible.


We see constant abuse from the Tencent ASN/associated ACE ASN, and I've memorized the china169 backbone asn as AS4837 because of thr nonstop crawlers splattered across their network ranges. It's not possible to ID the operator running the crawlers running from these networks but there's a clear signal of the origin of some of these entities.


Well, those companies have bots that identify themselves, and you can see what they're doing. Google especially have decades of experience of designing scrapers and seem to be able to the job of scraping the whole internet without causing problems in that time.

So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.

(It is worth pointing out that most of this traffic seems to be dumb: it's stuff like getting lost in generated link forests of some web apps or repeatly re-querying the same endpoint on a super-high frequency. This isn't exactly going to give a good return on investment for AI training data, especially since AIUI the main race for LLM performance now is in good quality training data)


>So, it's still possible they are also running a bunch of poorly-coded scraper bots through residential proxies, but it seems a little non-obvious why they would be doing this.

why "non obvious"? The easiest explanation is gathering training data sets, and is mostly caused by AI companies guarding their pile of essentially stolen IP (given how little they care about copyright) from eachother, without sharing any competition need to get their own and re-crawl to update it too


Well, why run both a well-behaved, easily identifiable bot that probably already gets them all the data they need (they seem to be spending most of their time and money on getting higher quality data than your average internet scrape), and this crap? Like, it's possible the obvious bots are a smokescreen, don't actually work well enough, and they are actually also reliant on data sources that contain 2 million copies of gentoo's bugs database. It's even possible that they are indirectly responsible for it, by buying datasets from shady sources, but again this requires a few jumps that I would like to see justified by evidence.


Most of the nonsense comes through botnets/residential proxies, so it's very hard to say.


I sometimes still wonder exactly who these scrapers are

This has zero to do with AI (it's just a convenient scapegoat that fits the narrative). Absolutely everything to do with those pushing for centralised control of the Internet.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: