That benchmark could really be a good AGI test. Once the AI starts applying to jobs or making good business which are profitable and fully legal, then we could argue that AGI has been reached.
If you could construct a sandbox to test this, where it doesn’t touch the real economy, then yeah it’s a great benchmark.
As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.
This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.