Your agent can write the test and run the test. Until now, it couldn't see the page.
Test failures will happen. Flake will happen. Writing new tests against a product deadline will happen. We've all been there and will be there again in the future.
Software development is predictable that way, and completely fluid at the same time, which is what makes every day a slightly different problem than the one before it.
As we've all started folding AI into that work, we've learned it's extremely good at producing a lot of code, fast. And one rule has held up better than any other: the better your inputs and the more context the tool has, the better the output. That's the whole reason we encode conventions as Skills, point agents at tickets through MCP, and hand them our patterns instead of hoping they guess right.
But there's a ceiling, and it shows up the moment the agent has to work against a running app.






