Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Good question.

When BF Skinner used to train his pigeons, he’d initially reinforce any tiny movement that at least went in the right direction. For the exact reasons you mentioned.

For example, instead of waiting for the pigeon to peck the lever directly (which it might not do for many hours), he’d give reinforcement if the pigeon so much as turned its head towards the lever. Over time, he’d raise the bar. Until, eventually, only clear lever pecks would receive reinforcement.

I don’t know if they’re doing something like that here. But it would be smart.



they’re not doing anything like that and you are actually describing the failed research direction a lot of the frontier labs (esp Google) were doing


Since intermediate steps of reasoning are hard to verify they only award final results. Yet that produces enough signal to produce more productive reasoning over time. In a way when pigeons are virtual one can afford to have a lot more of them.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: