In the past I managed to get measureable performance optimizing a harness by looking at few traces to see if the traces contained surprised, a lot of text in order to figure out how to use my custom tool, then renamed the tool, changed some parameters and it was already great across around 20 eval tasks in rust/typescript, I repeated the same more recently but I used an llm to look at the traces... didn't achieve the desired result, mostly due to how cost-prohibitive it's for me to run expensive models.