After three weeks of running Qwen2.5-32B on a DGX Spark, the number that surprised me most wasn't the throughput or latency. It was zero.

Zero structural errors across 2,859 code generation tests.

What I Tested

EvalScope with code generation tasks covering:

Structured JSON output