AI coding agents are getting better at writing code.

But I think we are approaching a more difficult question:

How do we know that an AI agent made the right engineering decision for the current state of a software system?

Passing tests is important.

But passing tests alone does not necessarily tell us whether an agent understood: