Periodically I try LLMs to see if they actually are useful for programming. This week I was writing an AsciiDoctor plugin. I’m not particularly familiar with Ruby (I struggle to imagine the mindset that believes Smalltalk, or, indeed, any language, can be improved by adding Perl syntax) and I find the internals of AsciiDoctor to be very poorly documented and hard to follow. So this seemed like a great task for LLM assistance: I know the problem domain very well (I’ve written other document processing tools) but not the details of the codebase I’m working on.
GPT-5.6 repeatedly sent me down misleading paths and wasted time, and gave suggestions that looked right but were completely wrong. It kept going back to the same thing that would cause a later phase to infinite loop. It eventually gave me a solution that was so convoluted that I couldn’t believe AsciiDoctor wouldn’t have a better way of doing it (I didn’t try it: if that was the only way of making it work I’d rather ditch AsciiDoctor than apply it, so I have no idea if it actually worked).
I actually solved the problem by finding a bug report from someone trying to solve a similar problem, which linked to their eventual solution. Reading that, I was able to find the missing conceptual bit in my understanding of how AsciiDoctor worked. And I finished the thing.
In total, the time spent with an LLM did not advance me towards the solution at all. None of the code that it suggested that I write ended up in the final plugin.
If I had used an ‘agentic’ setup, where I had given it my document, the desired output, and told it to keep trying until it did it, it would probably have got something that worked. The last throw looked like it would have ended up working, but it would have been an order of magnitude more complexity than I actually needed. And would have been completely unmaintainable: if all of the surrounding justification text had been comments, it would have looked like entirely plausible code, but with deeply fragile connections between the different phases.
Importantly, if I hadn’t solved it myself, I would not have really understood the gap between the complexity of a good solution and the LLM-generated one. I might have accepted it was the best available option, in spite of it being terrible. And this was fairly low-stakes code (processing a document for internal consumption), so that might have been fine. But doing that for anything used in production would be an absolute long-term disaster: you’d acquire all of that technical debt without knowing that it is technical debt, you’d just think the feature was done.
@david_chisnall I love how "agentic" is basically simulated annealing.