Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I concur, treating models as question and answer machines and judging them on recall is meaningless, unless you're measuring quantisation impact on a foundation model maybe.


100% agree, it's really just for kicks&giggles. The fact that the model answered correctly, unlike every other model of its size before it, still pleasantly surprised me.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: