You tweak the prompt. Run it against the three examples you always use to sanity-check. It looks better. Ship it.
That's not evaluation. That's vibes with extra steps — and it's the exact habit interviewers are trained to catch with one question.
This is one of the fastest ways an interviewer separates "I iterate on prompts" from "I've actually shipped prompt changes to production." Everyone iterates. Almost nobody evaluates.
How to answer it:
Name the trap first, out loud: eyeballing a handful of hand-picked examples is confirmation bias, not evaluation. You'll always find a few cases where the new version looks better — that's not evidence, that's cherry-picking with good intentions.







