Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

In a PDF or normal program there are clear semantics what is text and what is something else. On a canvas, everything is just made up of pixels. You'd need OCR Software to detect what is what and they won't ever be 100% correct unless you use only text and fonts which are made to be recognized by OCR Software.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: