Looks like most things definitely stump even a "smart LLM"... Best score on this is 30%. Which is what you should assume for tasks you give an LLM if they aren't exactly the same as an existing benchmarked task. They're just not that good for the purposes people seem to think they are. Very limited application space.