The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.
Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.
My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here
https://github.com/ed-is-ai/featherbench
Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.
My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench
Encourage everyone to eval like the devil