I agree with you, it could surely expand this answer (I am the person who wrote this little answer). Since this was originally just an answer to a question I answered via email (if I remember correctly), I didn't go into too much depth regarding ConvNets, because I just wanted to answer this question "briefly" in this mail. But if there's a demand for that (if it's useful to others) I may end up writing a tutorial on ConvNets one day :). However, there's already an excellent one out there, I highly recommend Dumoulin & Visin's "A guide to convolution arithmetic for deep learning" at https://arxiv.org/abs/1603.07285
I felt this could have been expanded on.. still not sure how the sliding windows map into units.