Retort is a framework for comparing coding stacks, with results for versions of Claude and local tests on a 64GB M5Pro for many languages. It was developed with Claude, and I was getting confused by all the components of the harness and stack, so I asked Claude to explain what is out there and how they are being tested. That's what follows:
Most benchmarks answer "which model is best?" Retort insists that's the wrong
unit. A coding result is produced by a whole stack — and the model is only
one layer of it:
language × model × weights-format × serving engine × agent/harness × context engine × sampling × prompt







