Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.
The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.