On the ScreenSpot-Pro benchmark, all effort levels achieve roughly the same score. I wonder if that is just a limitation of the benchmark, or if the effort levels actually do not make a difference for purely visual tasks.
> ScreenSpot-Pro tests whether models can locate the correct interface element in high-resolution screenshots of professional software.
> ScreenSpot-Pro tests whether models can locate the correct interface element in high-resolution screenshots of professional software.