Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.


Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.


The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: