TL;DR
Separating the generator from the evaluator improves quality and reduces premature self-validation.
The loop works best when feedback is explicit and based on clear rubrics, especially for subjective or complex tasks.
It is useful when the task has high value; for simple or easily testable tasks, it can become overengineering.
How to Separate Production and Evaluation in Tasks Without Ground Truth







