Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I am the only one who thinks google reCAPTCHA is just a tool that Google uses to train Machine Learning Algorithms? First it was used to help Google learn how to Read, now its learning to detect object, landscapes, ...


I'm pretty sure this is the advertised purpose of it. It stops bots AND helps machine learning.


*AND helps a company make profit by training their proprietary models.

If they were open, it would be a lot better (considering all the training is done by volunteers, too)


uhh. Everything they do is arguably for profit. That's why Google is an LLC. What do you expect them to do, put an "* this is done to earn us money" disclaimer on everything?


This is a little pedantic and doesn't change your point, but Google is not an LLC. It's a C Corporation.


I like to misidentify things to give them bad data.


Well, I do the same. I won’t destroy any of their dataset, or even have a measurable impact – but in turn, I won’t have a measurable impact in positive direction either.

If they want to make a profit from me, they should pay me, or they can license correct CAPTCHA results from me under GPL.


This argument doesn't really make sense to me. What browser, OS and device are you using to post this comment? Someone made profit when you bought the device, bought the OS and/or browser.

A healthy system needs some kind of motivation. In economies, that is profits/money. What's wrong with that? (I know I am being simplistic here but...)


> A healthy system needs some kind of motivation. In economies, that is profits/money. What's wrong with that? (I know I am being simplistic here but...)

Simple.

I buy a thing, I get something – in that moment the contract is over.

VW doesn’t come to me every 3 days with "You bought a car, to continue using it, take this and drive to Hanover and deliver it there".

When I bought my computer, or its parts, I bought them, I put them together, and that’s it. The manufacturer has never asked me to do work for them, or pay for them again.

I use ARCH Linux and Firefox – projects done by volunteers, and they profit by having a better product for themselves and others.

Google profits from ads, and from selling data. That’s a tradeoff, and a reason for me to try to use as few Google products as possible.

But when government organizations use ReCaptcha, and I have to work without pay for Google, then I have no choice, and it is not something I agree with.


The devices that I buy generally don't ask me to do work for them. If they do, for example by spying on my behavior, then I'm not so happy. On the other hand, I may accept some spying/doing work for companies if I get something in return - which happens for some Google products, like search.

When doing a captcha, I don't really get anything. It's something I have to do because the website I'm using has the problem that they can't find a better way to identify bots. So I do a captcha, fine. But, if there's benefits for whatever company offers them, i.e. I'm doing work for them without getting anything in return, or without everybody getting anything in return (the GPL option), then I'm again not so happy.


Napster, torrents, and the free internet warped an entire generation's perspective on things. This sort of "everything should be free" attitude isn't going away any time soon.


The advertised purpose was to digitize the worlds books and help libraries, digital humanities and the world.

This ended, and all trace of public good was erased (Check archive.org if you dont believe me) when Google execs mandated that only things that made money can be supported.


You do realize that it was the US court system that put the library-of-the-world version of Google Books on ice, correct? Larry & Sergey wanted a hippy "share all the books to everyone" view of this. The court said no, do what the book publishers say.


While you're mentioning the Internet Archive, you might also recall that our advertised purpose is to digitize all the world's knowledge :-)

I just had a chat with our book guy today about OCR corrections using captchas, one of our volunteers suggested it. We expect to scan 500,000-700,000 books this year.


The little captcha box does some processing before deciding what to show you. If you look suspicious, it can show you a more complex challenge. If you look like a normal browser, it might show you a house number to read (to improve its Maps product maybe). If you already have a cookie set because you already proved you're human, maybe it's just that check-box that says "I'm a human".


Unless you're able to explicitly state how this is done and/or how to trick it in think you're "high risk" - then seems like speculation; yes, I'm aware Google's said this, but never seen an proof of it. I've found bug in the past in the system that were easy to fix, told Google, but the bugs never were fixed.


Try doing a Google search from tor. The catches that tor users get are borderline impossible because of the suspicious ip and/or browser.


Not impossible, very often I get the "Body of water" one, but I do seem to get a challenge every 10 minutes or so, in each new tab.


I've experienced this first-hand. I used to always get the checkbox, but then one day I had to download a large series of related files from a website that used reCAPTCHA. I did all the clicking manually, but in a very repetitive and bot-like fashion, by opening a series of 10 new tabs at a time and then performing the same series of clicks on each tab to get to the download link. After a few minutes, I stopped getting checkboxes and started getting increasingly more difficult CAPTCHA challenges.


It's not speculation. It's section 2 of the paper, titled "Analyzing Risk Analysis System."


I can't find the original source, but the system for deciding whether to just show a check box or not was super complicated.Back around a year and a half ago when it was first released, some people on 4chan's technology board (/g/) were frustrated because they'd repeatedly be marked as suspicious (and never got the check box only) when posting on 4chan (which requires a captcha to be solved for each post). One user in particular was reverse engineering it and published a ton of super interesting stuff on Github (the levels of obfuscation were insane), but later got a job offer from Google (allegedly) in exchange for deleting it.


No, that's definitely what it is. It just also happens to be useful for blocking bots.


I agree with you




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: