TL;DR

Separating the generator from the evaluator improves quality and reduces premature self-validation.

The loop works best when feedback is explicit and based on clear rubrics, especially for subjective or complex tasks.

It is useful when the task has high value; for simple or easily testable tasks, it can become overengineering.

How to Separate Production and Evaluation in Tasks Without Ground Truth