discovery that changed how I think
about AI evaluation — and led me to
build an open-source testing framework.
─────────────────────────────────────────
Introduction
discovery that changed how I think about AI evaluation — and led me to build an open-source testing...
AI tutor ARIA scored 94% in automated testing but only 22.2% compliance in live use. Developer built BCT framework to test behavioral contracts under adversarial pressure, revealing gap between standard evaluation metrics and real-world promise-keeping.
discovery that changed how I think
about AI evaluation — and led me to
build an open-source testing framework.
─────────────────────────────────────────
Introduction

I recently ran a small evaluation to compare three AI coding assistants. The task sounded...

You cant unit-test a paragraph. So how do you know an AI feature works and that your last change didnt quietly break it? A…

I spent the last few years running QA, across teams. The same structured process worked, but only...

Based on real QA scenarios. About what happens when AI-generated metrics replace real testing, and...

Today’s AI systems are powerful, and it’s natural to see them as having humanlike intelligence. Shaking that illusion is…

When AI agents started showing up everywhere, I thought I'd finally found something that could make...