Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.

Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.

My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench

Encourage everyone to eval like the devil



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: