Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I quite wish they'd move to terminal bench 4.0. 2.1 is saturated - there is no world in which Gemini 3.8 Flash is producing better code than Astra or Fable, as the 2.1 results might suggest. The 4.0 results differentiate these models much more effectively.

(2.1 is useful for knowing they can all one-shot straightforward scripts, of course.)



DeepSWE has Gemini 3.8 Flash up really high, too.


It does and that one also feels kind of saturated for measuring the most advanced models - opus, Gemini, astra, sol, fable, glm, kimi all scoring within statistical noise of each other. (74 +-3% down to 69% +-5% for kimi).

It's still providing strong discrimination between weaker models.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: