In a PDF or normal program there are clear semantics what is text and what is something else.
On a canvas, everything is just made up of pixels. You'd need OCR Software to detect what is what and they won't ever be 100% correct unless you use only text and fonts which are made to be recognized by OCR Software.