There’s a good eval floating around somewhere and tl;dr they’re awesome but the benchmarks are cooked, you’re better off with Qwen 8B Q4 than 27B 1b or ternary.
Thanks for being skeptical, I maintain a llama.cpp-based client and it’s frustrating how high expectations are for local AI bc the median effort level means people mostly assemble their expectations and understanding via marketing soundbites
Thanks for being skeptical, I maintain a llama.cpp-based client and it’s frustrating how high expectations are for local AI bc the median effort level means people mostly assemble their expectations and understanding via marketing soundbites