16 local model configurations, 56 hidden-test coding tasks, 36 full runs, one 16 GB card

Every number below was recounted from the committed SCORES-*.tsv files, not transcribed from notes. The raw data - 36 rows, one per run, each with its full failure list - is RESULTS-q56.csv.

gemma-4-e4b is 7.5B parameters in a 4.97 GiB file. It scored 42/56.

devstral-small-2-24b is 24B in 11.90 GiB. It scored 40/56. gpt-oss-20b scored 38. Both

gemma-4-12b variants - one of them at Q8_0, more than twice the precision - landed at 42.5 and 41.