Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

GLM-5.3 is further proof that all >1T models are currently undertrained. I was looking at inteligence density ( https://www.pasteboard.co/6q2-5f92mtj9.png ) from recent open models (where parameters sizes are known) and taking DS-v4-flash as upper limit GLM-5.x can 3x its performance.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: