I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?
Might be cool to try different colors for the different attention heads instead of summing them across all attention heads, and making the backgrounds composed of stacked color layers? So you can see how each attention head attends to tokens individually.
Point taken, but there is a much more fundamental issue with it - and precisely why I wrote "very conservative".
It is a different problem if we pick two sets from the same data distribution, A and B, and first we have a score on A, then on B.
Here we re-run on precisely the same set of Terminal Bench 2.1 problems.
It may be that results are so random between runs that each single task has the same probability in a Bernoulli distribution.
But more likely, many problems are easy (i.e. each run will solve them consistently), many are too hard (i.e. no run is going to solve them) and only a fraction is somehow in between.
Maybe there is some good trick to find a proper distribution, but to my knowledge, we would need to run it at least two times on TB2.1 to get any more educated estimates.
That said, I am open to new ideas.
The main problem here is that a model that wildly fluctuates with 60% - 100% - 80% results will have the same wilson score as one that repeatedly scores 80% - 80% - 80%. So the 'confidence interval' bar is meaningless.
I'm not that well versed in statistics, but a standard box plot is probably the best alternative
I wrote this blog post myself, with AI for proofreading (typos and grammar, but not style). There were a few singular sentences for which I had a writer's block, but not much besides that.
So, if there are irrelevant remarks, these are mine. :)
Charts are vibe-coded - but it took quite a bit of hand-holding to get something decent. And the logarithmic scale for model size is my conscious choice (against Claude's initial ideas).
> The phase of free fall is predicted in a purely quantum manner to have a dependence m/6 g^2T^3/ℏ + gmzT on the free-fall time T, where m is the mass of the object, g is the gravitational acceleration relative to the surface of Earth, and z in the spatial coordinate in the direction of gravity. This
prediction follows the calculated phase accumulated by an object accelerating in a linear potential, and has been made starting from almost one hundred years ago by Darwin, Kennard and others.
Apparently the phase shift is derivable from just adding a linear potential term mgz to the Hamiltonian.
It contrasts with times before, when we actually had to wait for a monthly magazine, and even if we wanted to watch a movie, it was aired (say) the next Tue 7PM.
While now we have all convenience, there is no anticipation - or rest.
> Intelligence Index vs. Cost per Intelligence Index Task from Artificial Analysis. Note that it is based on score of benchmarks like Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond - not necessarily fluid intelligence like in abstract puzzle games of ARC-AGI-3 or Baba is You.
On wikitext2 there was no difference. I concluded these have no long-term dependency and I should use Linux kernel. Still, the same.
So yes, KLD depends on the dataset. Still, it does not measure what any e2e test does.
reply