But now that 3Taps has found an ingenious way to get at the data with zero extra bandwidth cost to Craigslist (by retrieving it from the Google cache rather than CL itself), it's clear that what Craigslist really dislikes is competition.
While Craigslist is probably within their legal rights here, this case shows that for all their talk about their benevolent aims, Craigslist is no different from other companies.
When I think of Padmapper, I think of the naughtiness:
"Though the most successful founders are usually good people, they tend to have a piratical gleam in their eye. They're not Goody Two-Shoes type good. Morally, they care about getting the big questions right, but not about observing proprieties. That's why I'd use the word naughty rather than evil. They delight in breaking rules, but not rules that matter. This quality may be redundant though; it may be implied by imagination.
Sam Altman of Loopt is one of the most successful alumni, so we asked him what question we could put on the Y Combinator application that would help us discover more people like him. He said to ask about a time when they'd hacked something to their advantage—hacked in the sense of beating the system, not breaking into computers. It has become one of the questions we pay most attention to when judging applications."
Naughtiness was using CL to seed their site in the first place. After the Cease and Desist, using 3Taps as a source CL data was simply walking into a lawsuit.
Generally, I don't think you want your naughtiness to tie you up in a legal battle.
Case will never go to trial. Maybe a "legal battle" was the only option.
Can CL "win"? Can they stop everyone else, PadMapper is but the first of many, no doubt, from re-displaying facts (classified ads) in different formats?
Craigslist is probably within their legal rights here
I would really call this into question for the copyright claims. The key claim is copyright infringement, and the listings on Craigslist are almost certainly unprotected, much like telephone directory listings. They are statements of fact rather than creative works.
Craigslist may attempt to claim copyright over reproduction of their database as a whole (compilation), but the Supreme Court has ruled that in order to be eligible the compilation must be "original in its selection, coordination, and arrangement", and Craigslist can hardly claim that.
The breach of contract claims may be stronger though. The pages on the Google cache presumably still contain the ToS, and if they apply (who knows?) then PadMapper's use of Craigslist would likely constitute a breach. A ruling on this would be very interesting.
Another thought is that padmapper is not /copying/ the data - They are summarizing it.
So there could be a entire listing for a SF apartment with lots of details and padmapper would not copy any of it- instead they would just make notes - 3 bedroom, location, price, pictures and then link back to CL if someone clicks the listing on PM
>The breach of contract claims may be stronger though. The pages on the Google cache presumably still contain the ToS
Would this work? Can a phonebook just put a section at the beginning of this page that by reading this book you are agreeing to the ToS which state you can't copy it?
I know most ToS contain a section about by using this service you are agreeing to the ToS.
But I think there is a good argument that Padmapper isn't actual using the service since they aren't getting the information from Craigslist's servers.
Am I using a service, and thus bound by contract if I look at a screengrab of a website on a third party website (that contains a ToS).
I've said before though that Craigslist could invent fictitious entries, and sue Padmapper for copying those creative works, just like mapmakers do with fake towns.
Would it work? Well, web pages are cached by ISPs all the time. The nature of the internet is that data is buffered all over the place. The argument seems to be that because the data has been at rest for a few days, the ToS no longer applies - this seems like a fairly weak argument, but who knows.
As for phone books, a shrink-wrap license on a CD-ROM phone book which prohibited copying was upheld by the US courts (ProCD), so that copying was a breach of contract even though the underlying data was unprotected.
> PadMapper isn't actually 'using' the service.
Using is a wonderfully subjective word, and I'd expect accesing, and making use of to be acceptable synonyms.
> I've said before though that Craigslist could invent fictitious entries, and sue Padmapper for copying those creative works, just like mapmakers do with fake towns.
These are known as "trap streets" and US federal court has ruled that they are not protectable. Map makers are able to sue because the rest of their map is protectable; the trap streets simply catch the infringer red-handed.
Hm, would you not say that what Craigslist really dislikes is their competition piggy-backing off their data? Craigslist seemed quite happy with the existing situation until Padmapper launched their own listing service.
Though I agree that it's not in the spirit of a .org, if such concept exists.
"You automatically grant and assign to CL, and you represent and warrant that you have the right to grant and assign to CL, a perpetual, irrevocable, unlimited, fully paid, fully sub-licensable (through multiple tiers), worldwide license to copy, perform, display, distribute, prepare derivative works from (including, without limitation, incorporating into other works) and otherwise use any content that you post. You also expressly grant and assign to CL all rights and causes of action to prohibit and enforce against any unauthorized copying, performance, display, distribution, use or exploitation of, or creation of derivative works from, any content that you post (including but not limited to any unauthorized downloading, extraction, harvesting, collection or aggregation of content that you post)."
> You also expressly grant and assign to CL all rights and causes of action to prohibit and enforce against any unauthorized copying, performance, display, distribution...
Rightshaven (copyright troll) was granted the same right by the copyright holders they represent, yet the judge ruled they didn't have standing to sue on the copyright holder's behalf.
The Righthaven case is slightly different, though I hope the logic still applies. (I'm not a lawyer, but I read the Righthaven opinion[1] when the Padmapper/3Taps workaround was originally discussed.) The difference is that Righthaven was granted merely the right to sue on behalf of the original copyright holder, but none of the exclusive rights that copyrights actually bestow upon their owners. It could be interpreted that posters grant Craigslist some of their exclusive rights ("copy, perform, display," etc.), and suing to protect those rights may be legally kosher. I know of no such precedent.
My personal interpretation/hope is that the right to sue for copyright infringement is nontransferable, which would give Craigslist no standing to sue. Individual posters could sue, however.
That's a given. The only way they'd have authorization by the owners of the data would be to e-mail the poster of every CL listing and ask for it. We know they don't do that. The key point is whether there's copyright infringement at all, not whether it was authorized.
And you automatically grant Craigslist your first born child and all future earnings and your left index finger. Laywers love to fill these things with unenforceable outlandish crap. The courts are more discerning.
Err, that's actually a very reasonable Terms of Service.
It says that you own everything you post, but you are granting them the rights to use it. The only unusual bit is that you are additionally granting them the right to go after people who scrape your content on your behalf.
That's not that unusual for sites that contain user-generated content. If someone's hosting a blog full of scraped YouTube videos, YouTube can go after them without contacting the owner of each individual video.
The grant is not exclusive and the authority to decide what is an is not authorized is not claimed by CL. If it is unauthorized they claim rights and causes. You can't have a non-exclusive grant without this distinction.
Exactly right. And if you want your data to be reposted on another site, then you should repost it.
Craigslist is delivering exactly what it its users signed up for (no more no less). People posting ads on Craigslist do not necessarily want or intend for it to be reposted on other sites.
"People posting ads on Craigslist do not necessarily want or intend for it to be reposted on other sites."
That argument appears to be invalidated by Craiglist's own terms of use, which say that when you upload a listing to Craigslist they can syndicate it wherever they want (http://news.ycombinator.com/item?id=4287519).
Well, that's a fair point, but I don't think that right matters much unless and until they exercise it. A better argument (against myself) is that they're apparently willing to sell some type of access to your posts for use on mobile apps.
Still, I think there's something admirable in the simplicity and transparency of interacting with Craigslist. What you see is pretty much exactly what you get.
The key point is that the TOS says that you grant Craigslist the right to redistribute your listing where ever they see fit, it doesn't say some third party entity has the right to do that.
As far as I'm aware, if you post information in a place where it will be publicly accessible, then you are tacitly agreeing that third parties will be able to access it and use it as they please. That's not a right that needs to be given to these third parties. It exists from the get-go, and must be explicitly taken away by something like copyright.
So I can legally start amazon-copied-reviews.com and scrape every product review from amazon.com with my own referrer links to Walmart without fear of repercussion? Sweet.
> And if you want your data to be reposted on another site, then you should repost it.
I had to giggle a little bit at this in the grand scheme of the internet, sorry.
When I posted something on Facebook Marketplace, I started getting emails and comments from other "market" sites that Facebook had cross-posted my listing to.
While it was annoying to not know this up front, the fact that it was more visible and getting more bites because of it only helped me make the sale quicker. If a service wants to piggyback off of another to make my postings more buoyant, as a user and seller, I don't have a problem with it.
In response to your "only helped": on one of the other Padmapper stories here on HN, someone wrote that after their item sold and they cancelled the Craigslist ad, they continued to get contacted about the item because third parties had scraped the ad and did not stop displaying their copy when the ad disappeared from Craigslist.
This happened to me on the Facebook ad too, but I preferred it it over the lack of any response at all - it proved the services were used. Easy to ignore them or send a generic reply.
Padmapper's isn't "reposting" any given listing -- it displays an abbreviated digest of the listing in its search results, and then if the user clicks, it takes them to the original listing.
If this is illegal or otherwise objectionable without an explicit agreement from each Craigslist poster, I'm not really sure how a search engine or even descriptive hyperlinking is kosher without explicit agreement from each website indexed or referred to.
The point is Google respects publishers' desire not to be indexed. Padmapper does not. Robots.txt is simply a common method for conveying that message. It's not like Padmapper could argue they didn't know CL was unhappy; they got a certified letter!
This is just semantics. Google respects publishers who do not want their sites listed.
True, but that's irrelevant to the comment I was replying to. We are talking about user expectations with regard to content they posted being made available on other sites.
Craig is doing things right: he is not greedy and he doing it with vision of general purpose. If craiglist is run by a person who investors like to choose as a founder they will be sold for 10M a long time ago.
The problem with a lot of startups in SV is that their only objective is to make money. Nothing else. Craigslist is very very rare exception.
Flogging yourself doesn't make you righteous. Doing good for the world does.
Craigslist wastes several human lifetimes worth of time every month through maintaining a monopoly product with a shitty UI and refusing to let anyone innovate on top of it. They hold back progress and they are evil. They are the IE6 of classified ads.
Not sure I would call that ingenious. But your point is sorely need of being made more often.
How many other websites use arguments like "bandwidth" to falsely portray competitors who access their publicly shared data as somehow in the wrong?
Many. Some here on HN. No need to name names.
No doubt even Google would complain about people "scraping" search results.
To me, it is a joke. Because the people who complain use automation to access, retrieve, organise and serve information and thereby establish their business. Only then to try to forbid others from using automation to do the same.
And all the while, it's NOT THEIR INFORMATION. This is not Craigslist's data. It's users' data.
It belongs to users, who are today's "publishers" and possess all those good ole publisher's rights. (Though they may naively license them out.)
The users chose to give Craiglist permission to use their data. They did not choose to give PadMapper their data, or else they would have posted their listing to PadMapper.
They did give Craiglist permission to prevent others from using their data without permission. In this context, scraping any version of Craiglist's site (whether CL itself or a third party cache) falls within Craiglist's rights under the license they were given, and within the user's expectations of what Craiglist will do with their data.
Weak argument. Did they give Google "permission" to access the data (and store and republish it)?
Google is allowed in robots.txt. But that is not exactly what I would call an agreement.
The simple fact is this info is on the public web which, by its nature, copies and transfers data. That's what the web does. You upload something and it goes "viral". You have principles like the "Streisand effect" to contend with.
This goes back a long way. No doubt judges remember. The Ken Starr report on Ms. Lewinsky. Some random classified ad. Like it or not, information gets desseminated.
If you want to protect and restrict access to data, then you do not upload it to the public web. You put it behind access controls, e.g., a password. This is common sense.
If anyone has a claim here, it's users who do not want their ads on PadMapper (if there are any). CL has no standing and their motives are both pathetic and transparent.
Craigslist is a business. What makes it right for another company to profit from the contents they generated through the platform that they build?
Yes, they got an early advantage into the market and has the critical mass that many company can seem to compete with but why take that away from them because they refuse to update/add new features.
Instead of piggybacking on them and relying on their data to earn money, shouldn't Padmapper focus on building their own content. Isn't that where innovation comes from? Beating an existing company by creating a better platform?
Craigslist is a business. And a monopoly, like Microsoft in the 90s. Didn't they have the right to decide which browsers are packaged with their OS? Don't they even have the right to decide which browsers run on their OS?
Exactly, they've pretty much run every newspaper classified out of business. Consumers have no choice but to do business with craigslist if they want to list something in a classified.
That's why website owners should focus on creating another classified website that can compete with Craigslist instead of relying on their contents. That's where innovation starts.
And its not about a matter of choice the users have. There are many competitors in this market, and yes Craigslist dominates every one of them because they had an early advantage on the internet. Small sites can't just leech off the contents on their website and slap ads on it to make money.
Lets say Craigslist was a print company that produce and distribute classified as. Will it be right for a small company to steal their content and slap their ads on it and distribute it themselves?
Your argument won't work here, because there are many major newspaper that do 1000x in revenue and distribution than independent newspapers.
Just because they have competitors doesn't make those competitors viable. Apple and Linux still existed when Microsoft was prosecuted by the DOJ.
The vast majority of people searching classified ads are searching craigslist, therefore if you're trying to list something in a classified ad, you're forced to use craigslist.
Sure you could use another service, but Netscape could have also just sold browsers only to Linux customers. It's all about the numbers.
But Microsoft largely escaped the DOJ antitrust case because of competition with Apple. (MS invested in Apple to retain a competitor.)
People aren't really forced to use craigslist. That the vast majority of people choose to search CL isn't good enough. Another company could spend whatever it takes to get people to search their classifieds instead. As long as CL can't or doesn't block that (in contrast to stuff MS was doing that started the DOJ case against them), the competition is viable.
I'm sure Craig was speaking truthfully when he made his comment. But he is but one piece of a whole, and that whole includes folks who count the pennies. And business person associated with Craigslist would know that the only value Craigslist has is its listings. Giving those away to be re-used by another web site would never fly with such a person. I don't doubt for a minute when you tell someone is taking money out of your pocket that you need to fund your operations (I know, I know its more nuanced than that but its the gist of the argument) well you an roll over or you can fight.
If they go to court and get an opinion it should be really helpful as a guide for other entrepreneurs who see re-processing the information on the web in new ways as the foundation for their business.
If they go to court and get an opinion it should be really helpful as a guide for other entrepreneurs who see re-processing the information on the web in new ways as the foundation for their business.
You mean, like Blekko?
The simple, inconvenient truth is most folks who are making money from the web, like search engines, are not content creators (nor content owners), they are content publishers... who publish for free. "Are you a non-technical person who wants to get something onto the web? No problem. We'll help you with that, for free. Just give us some personal info about you so we can solicit money from advertisers."
(Placement, e.g., paid placement, where the eyeballs are more likely to see something, for a fee, is another matter.)
In the case of search engines there is already a lot of case law from people suing Google of course. As with most things its a spectrum. Using a search engine as an example (and disclaimer I work for Blekko, a search engine) a search engine crawls the web, then it computes a number of parameters about the page its crawled (what its about, how many people link to it, who does it link to, Etc.) and creates a new piece of information called a 'rank'. Then when a search query comes in the query is used to create a way of recognizing a 'target' page and then the rank is used (and in our case slashtags too) to decide what pages you might be looking for. The results, also include snippets from the page to help the user evaluate whether or not the page is the one they want.
Now that use has generally not been highly contested, people want their pages to be found and so they tolerate search engines searching them. They can be explicit in what pages they want searched and which they don't using robots.txt. So that relationship is pretty well understood. People who ban Blekko (and presumably anyone else) from their robots.txt file are not crawled by us, we recognize and honor that it is there choice if they want to be in our index or not. On the copyright issue however it has been pretty clearly established that 'page rank', like someone's review rating on a movie or an application, constitutes an original work of the creator. There is a lot of experience with things like book reviews where the review, using snippets to illustrate the review, and a rating, are both fair use and the original work of the reviewer.
Google however got in trouble with their news aggregation service. And the bulk of much of the arguments there, were that the snippets were so complete on the news page as to exceed 'fair use' exemptions, and that by aggregating these pages they were 'stealing' traffic that might otherwise go to the news site. The results on those cases were mixed, with some newspapers being removed from Google's index, and others not. Generally everyone that was removed has since been replaced (at the request of the news source) because Google does drive more traffic to a web site than any other web service. So in this case while Google was found to violate the copyright of these news organizations by indexing their newspapers without their consent, the papers later found it in their best interest to give their consent.
Craigslist and Amazon and Ebay are a third kind of question. They are a collection of 'facts' (as many have pointed out) which are derived by a process (placing ads). And in the 'old' world the courts have generally sided with the person who had paid the economic cost for creating those collections. And as PadMapper and others before them have shown, is that there is a great temptation to use those same facts and re-package them into a new collection. This pretty naturally sets up a commercial tension between the original collector and the new user of those same facts. That seems to open another front in copyright litigation and policy. So if this court gets an opinion published it cannot help but be influential as there don't seem to be very many in this space. That could be because judges think the right answer is 'obvious' but I seriously doubt that to be the case.
I have heard people saying "we love competition". I wonder if they really mean it. Not saying one cant't love competition, just that the phrase is used so commonly that it's hard to digest.
I, for one, hate my competitors. Basic insticts perhaps?
Is not disliking your competitors something that is practiced by majority? OR is it at least very common? Common enough to point at a company that doesn't follow it?
It depends; If you are trying to start a new industry, you'd like your competitors more, because they validate the market. My university lecturer told me in the 90's (or 80's) he was starting a bottled water company; Back then no one drinks bottled water. Every month or so the half dozen bottle water company startups in Australia have a meet up; and they were all friends.
It's easy to interpret your post to mean that stealing/reusing data just because you can is fair competitive practice. Please do clarify because your words mean a lot to people and your post might be interpreted that way.
Hilarious to expose their sanctimonious nonsense though.
Can you comment on the status of Priceonomics (YC) in this context? They're building a rich interface on top of CL listings data, presumably using 3Taps to get the volume of data without drawing attention.
Is the endgame for these companies to replace the content source, or to hope (or fight for) a legal precedent to 'open' CL data for third party usage?
For the immense value they provide to the hundreds of millions of users, having a 100-200M revenue is not "swimming in [insane] $".
They're profitable alright, but let's not kid ourselves that they have by choice left a LOT of money on the table, which is why all these value-added services are trying to take a share of the CL-pie.
The AIM Group estimated $115M revenue for craigslist in 2011, after $141M in 2010. I don't know anything about the AIM group, but the report is referenced in the NYT (presumably they vetted the source, but...). http://aimgroup.com/2011/10/06/craigslist-revenue-falls-off-...
http://www.quora.com/Why-hasnt-anyone-built-any-products-on-...
But now that 3Taps has found an ingenious way to get at the data with zero extra bandwidth cost to Craigslist (by retrieving it from the Google cache rather than CL itself), it's clear that what Craigslist really dislikes is competition.
While Craigslist is probably within their legal rights here, this case shows that for all their talk about their benevolent aims, Craigslist is no different from other companies.