Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I clicked through and it showed Qwen at the top at 55.4 compared to 55.3 for Opus Max. I have a screenshot.

Then I clicked away and back, and now it goes Qwen second, with 58.4, to Opus Max at top with 59.2.

I have screenshots of both. The description above the chart is the same in boh cases:

> Artificial Analysis Agentic Index > Represents the weighted average of agentic capabilities benchmarks in the Artificial Analysis Intelligence Index (GDPval-AA v2, ³-Banking)

What happened? How can the scores change so much in a few seconds?




> HLE, AA-LCR and AA-Omniscience are now graded by GPT-5.6 Luna (medium), replacing GPT-4o, Qwen3 235B A22B 2507, and Gemini 3 Flash Preview respectively. These checks are now unified under a more capable modern model, selected for strong agreement with human judgment in our grader validation

Interesting that they chose a nano-sized model from OpenAI to be a grader for benchmarks involving knowledge and hallucination.


What's interesting is that if you ask 5.6 Sol or Opus 5 they will tell you it's a bad idea to have the reviewer be the dumber of the set as it can't judge them properly to decide who is right, and thus if one is better because it found an answer that's better but contradict the obvious it would be biased against. I know because I just had a consensus conversation with them this afternoon about a design that was similar (though about something completly different than judging agentic quality or whatever).


Fixed the result, eh? In both senses of the word.


Can someone please explain what changed, when it happened, and whether it was surreptitious?


I have been suspicious of these AI leaderboard sites for some time now, and this only increases that suspicion.


In that case they should clearly label that this is a new benchmark.


What was the change?


Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality checking/grader models, and brings the latest ³-Banking version to Artificial Analysis. Regular updates are normal for us to keep our benchmarks up to date.

The order changes but I think the story discussed in this thread holds - this is a very impressive release and Qwen3.8 Max is a huge step up in agentic capabilities.

Relevant blog post (also linked to by others): https://artificialanalysis.ai/articles/artificial-analysis-i...


You gotta admit the timing looks very suspicious.


Luna pricing was just cut by 80% https://www.eesel.ai/blog/gpt-5-6-pricing and as the blog post states is a more accurate judge than the previous methodology.


> You gotta admit the timing looks very suspicious.

Do you mean the timing looks like: "We're SV tech-bros. Our benchmarks showed a chinese model above what's considered the best model at the moment. So we quickly modified the benchmark so that our SV tech-bros don't look like they're losing to a chinese model"?

That's indeed a bit fishy.


They could just have avoided all of this by not publishing the benchmark until the new methodology update.


They should probably freeze the results before publishing.


Welp. That didn't last long


Same, they just updated it. Hacker news effect?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: