I had a system that looked ready.
One AI model designed the architecture.
Another implemented it.
Additional models reviewed the output, checked the logic, and produced validation documents.
Synthetic legal cases passed several internal checks.
I built a human-directed, no-code multi-model workflow with Google Drive. Synthetic tests passed. A real legal case exposed critical failures.
A multi-model AI system passed synthetic legal-document tests but failed on real data, misattributing claims and monetary values. Production AI requires real-world validation—inter-model consensus is insufficient to catch critical failure modes before deployment.
I had a system that looked ready.
One AI model designed the architecture.
Another implemented it.
Additional models reviewed the output, checked the logic, and produced validation documents.
Synthetic legal cases passed several internal checks.

A story about formal verification and adversarial testing. About systems that are mathematically...

Six months ago I started running every non-trivial piece of code through AI before it shipped. Not...

Six months ago I kept reading the same story. Developer uses Cursor or Claude Code to ship a feature....

Most legal AI tools focus on summarising documents or answering queries. While they’re useful, they...

I spent weeks auditing the smart contracts behind my Web3 banking project, SovereignBank Web3, the...

Six months ago I got fed up with my AI code review tool. Two problems. It charged roughly 3.5× the...