I built a GitHub App that auto-generates adversarial tests for AI-written code — here's how it works

Six months ago I kept reading the same story. Developer uses Cursor or Claude Code to ship a feature....

venerdì 19 giugno 2026 New tab

TL;DRAI

Khwand auto-generates adversarial tests for AI agents (Cursor, Claude Code), detecting semantic errors that CI misses. For managers deploying AI development, it automates edge-case detection and fix generation—reducing production-failure risk that standard QA can't catch.

390 words~2 min read

Six months ago I kept reading the same story. Developer uses Cursor or Claude Code to ship a feature. CI goes green. Merge lands. Three days later, production breaks in a way no test caught.

The failure mode isn't the model being wrong. It's that the tests being run were never designed for what an AI agent might do. The agent writes code that's syntactically correct, type-safe, and passes every existing check — but introduces a semantic error nobody scripted a test for.

So I built Khwand: a GitHub App that generates those tests automatically on every push.

How it works

When a push event hits the webhook, Khwand does four things:

I built a GitHub App that auto-generates adversarial tests for AI-written code — here's how it works

I built a GitHub App that auto-generates adversarial tests for AI-written code — here's how it works

Related reading

AI doesn't write bad code. It writes plausible code — so I tried to break my…

Reviewing AI-generated code before shipping: why I built Sego

I Added an AI Gate Before Every git push with no-mistakes 🛡️

The most dangerous line of code your AI agent writes is the test that passes

I Built a 'Production-Ready' AI Agent Framework. It Was a Lie. So I Fixed It.

The End of Manual QA: How I Built a Self-Testing App with Claude Code and…

Related reading

AI doesn't write bad code. It writes plausible code — so I tried to break my…

Reviewing AI-generated code before shipping: why I built Sego

I Added an AI Gate Before Every git push with no-mistakes 🛡️

The most dangerous line of code your AI agent writes is the test that passes

I Built a 'Production-Ready' AI Agent Framework. It Was a Lie. So I Fixed It.

The End of Manual QA: How I Built a Self-Testing App with Claude Code and…