Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Great reply. My question based on this comment,

>My advice is assuming you’d like to be a person that trains/deploys ML models to solve problems in industry. This is much different than an ML Engineer, who’s implementing algorithms in low level languages and squeezing out efficiency. Obviously that would require a much deeper understanding of SWE. And a totally different person is an academic researcher that’s developing theory or technique. It’ll be hard to do that without a PhD

Can one not only train/deploy ML models, but in addition to that be able to implement the algorithms in low level languages and also be able to develop theory?

I’d imagine these are all skill sets that someone in PhD program could pick up.

If they could do all three, what kind of job should they be looking for?



I think it’s unlikely to become expert in all of those things. If you do, it’s over the course of an entire career, not to get started. I guess it comes down to how much expertise is “enough” for you. Naturally, if you split your time across 3 domains you won’t be as expert as someone who dedicated all their time to going deep in one.

In the context of a big company, I think it makes sense to have a specialized workforce. Why look for the one in a billion person that can publish top quality theoretic papers and then implement them on distributed gpus in an optimal way while also building simple Random Forest models for your business? I’d rather that person do more of the most valuable thing, and then hire someone else to do the rest.


This answer makes sense.

I suppose my question is more along the lines of, if someone is specializing in deep learning in a PhD program then shouldn’t they at the very least be able to implement models and also know optimization tricks?

In other words shouldn’t they be able to develop enough skills to go deep in one area but also know enough to be dangerous in the other three domains?


I think I agree with you with the caveat that it would depend on what they're researching. If they're researching new model architectures, I don't think it makes sense for them to try to implement the algorithms from scratch in C++/CUDA to do distributed GPU training--why not just use TensorFlow? But if you're researching distributed tensor computation, then that's your bread and butter.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: