It’s much easier to get your AI agents to do something consistently than it is to get your human colleagues to do it. Heck, personally we know these things are the right things to do, but we are just too busy and our minds wander to what we think our more productive uses of our time. That completely changes when it’s agents being instructed to do the work instead.
The crazy thing is that it s looking like AI might be able to write software better than humans eventually simply because they do not get bored of doing tedious tasks. Many things that we know work but don’t practice because humans aren’t very scriptable are now viable and can be easily applied with agents. For example, I’m finding a lot of success in using separate agents to write implementation and tests from a common specification and then using an auditor to run the tests so neither agent is contaminated by the other’s work. This is just part of the clean room engineering process that was developed for people by IBM in the 1980s, and it was shown useful then, but with AI it can be widely and consistently applied.
They'd be 10x better than us already if task tedium was the problem. It's design sense that they're missing.
I find that in the areas where people think that LLMs excel at coding and don't like doing manually it's usually because the human was inclined to slop out repetitive boilerplate and thought that was the only way. Tests are usually like this, sadly.
LLMs are definitely good at providing reams of duct tape (which is drudge work) to patch up those bits of the code base where the code sucks. The problem is that duct tape is not the most architecturally sound construction material.
Plenty of languages are rich with boilerplate (think Java, C++...). Expressing intent succinctly is non-trivial, and especially so when you want to provide a rich vocabulary like modern languages do.
While modern tooling helps, the form of expression still means cognitive load (in reading code, deciding what to copy-paste and refactoring later).
And sometimes, tooling which is succinct brings a whole can of worms with it (think pytest with assertion rewriting and unexpected behaviour of your .pyc files — if you are familiar with Python as your nickname seems to suggest :)).
> They'd be 10x better than us already if task tedium was the problem... the areas where people think that LLMs excel at coding and don't like doing manually it's usually because the human was inclined to slop out repetitive boilerplate and thought that was the only way.
There is so much work out there that is repetitive / boilerplate / tedium. If you get to personally work on interesting work more than 50% of the time (pre-LLM) I'd say your job is #blessed.
IME if a technical task is tedious and repetitive it is nearly always because the system was badly designed or because it wasnt automated properly.
My job is automation and system design, so if I can't fix or work around these things that reflects poorly upon my skills.
Some people treat writing tests as inherently boring because their test frameworks usually suck. If they don't suck and a test is a very close approximation of a spec, it's not boring at all, especially if you use it as a means of codifying a spec before implementation.
Don't disagree that that's the system not working well.
You're lucky to be in a system that works out for you. Not everyone is in such a lucky position. LLMs help automate the tedium out of their job, and hopefully, have a bit more extra energy to make a better system within their immediate locus of control.
I'm sure you're where you are in some part due to hard work, skill, decisions. But surely you don't believe that it's purely down to skill where one lives/works at?
The crazy thing is that it s looking like AI might be able to write software better than humans eventually simply because they do not get bored of doing tedious tasks. Many things that we know work but don’t practice because humans aren’t very scriptable are now viable and can be easily applied with agents. For example, I’m finding a lot of success in using separate agents to write implementation and tests from a common specification and then using an auditor to run the tests so neither agent is contaminated by the other’s work. This is just part of the clean room engineering process that was developed for people by IBM in the 1980s, and it was shown useful then, but with AI it can be widely and consistently applied.