Skip to content

Random Thoughts on Agentic Software Engineering – 2026-09-30

Coming back to work on a codebase where I was part of the team that wrote the first lines of code: Defender for Storage Malware Scanning. I’m warming up by fixing a flaky test. Obviously, with #Copilot.

I’m still familiar with most of the codebase, the architectural decisions behind it, and, perhaps most importantly, many of its edge cases, even though I haven’t worked directly on the team for more than a year. This is a large system with multiple components and services, and high performance, latency, and availability requirements. There is a lot of context behind why things work the way they do.

I’ve used agents on other projects for a while now, but this experience felt different. In those projects, I wasn’t an expert in the codebase. I asked the agents a lot of questions, asked for second reviews, and looked at the resulting code changes myself (sooo 2025 of me 😁). Agents made mistakes, of course, but overall, the flow worked pretty well.

This time, the agent made changes that looked good. They passed the tests. They were validated by agentic reviews. I’m pretty sure that if I hadn’t been deeply familiar with the system, I would have approved them too.

But they were wrong.

And not obviously wrong. They were wrong in a subtle way that was difficult to detect without understanding some of the assumptions and edge cases behind the code being tested. I happened to know those assumptions because I wrote that code (and the flaky test 🤦‍♂️).

That made me rethink some of my experiences with agents on codebases where I wasn’t the expert. How many times did I see a change that looked good, passed its tests, survived another agent’s review, and conclude that everything was fine simply because I didn’t know what I didn’t know?

We’re getting very good at building systems where agents do the implementation, testing, and review, with less and less human involvement. And most of the time, that seems to work remarkably well. But maybe the interesting problem isn’t whether agents make mistakes. Of course they do. Humans do too. The question I’m now thinking about is: how do we detect the mistakes that require deep, accumulated system knowledge to even recognize as mistakes, especially when there isn’t a human expert in the loop who already knows where to look?

Still thinking about what this means for how I use agents in software development.

Has anyone run into something similar? More interestingly, have you found practices that help with this class of problem?

#AI #agents #agentic-software-development

Published inProgrammingThoughts

Be First to Comment

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Discover more from Musings of a Strange Loop

Subscribe now to keep reading and get access to the full archive.

Continue reading