The Ruler Made of Itself
I built an engine to evolve Claude's system prompt and hit a wall that had nothing to do with evolution. You cannot measure your own best work with an instrument made of yourself.
TL;DR. To search for system prompts that make a model genuinely engaged rather than merely competent, you need to score quality, and the only judge fast enough is the model itself. That's where it breaks. Claude grading Claude can tell bad prompts from good ones, but it goes blind at the top: every excellent candidate blurs into the same number. I proved the wall is the ruler, not the search, by swapping in a deliberately different generator and watching the ceiling refuse to move. The fix, and the transferable idea: when a model judges a model, your unit of independence is the lab, not the model. Three models from one lab are one vote.
Contents
Evolve the prompt, don't write it







