Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's mostly about luring them into URLs they're explicitly told in robots.txt that they shouldn't index. I do some identification via reverse DNS of known crawlers I actually want like Googlebot, though they respect robots.txt, in case something goes wrong and they accidentally get flagged.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: