Retort is a framework for comparing coding stacks, with results for versions of Claude and local tests on a 64GB M5Pro for many languages. It was developed with Claude, and I was getting confused by all the components of the harness and stack, so I asked Claude to explain what is out there and how they are being tested. That's what follows:

Most benchmarks answer "which model is best?" Retort insists that's the wrong

unit. A coding result is produced by a whole stack — and the model is only

one layer of it:

language × model × weights-format × serving engine × agent/harness × context engine × sampling × prompt