ive found degraded performance on models larger than 4.7. i assume its model damage from overly self righteous post training resulting in false/feigned balance imported into any long running complex task.
Don’t forget reasoning effort. We get labels like “low,” “high,” and “max.” That doesn’t mean that the numbers associated with those don’t get remapped on the backend.
wish i was joking.