Originally published on my Substack. I'm a Microsoft MVP based in Japan, writing in English about the AI agent systems I actually run in production.
Local AI models keep multiplying. But comparing numbers on model cards alone doesn't tell you which one to actually use.
Does a higher parameter count mean smarter? Does MoE mean faster? If a model is popular on AI Arena, is it good for my own work? Each question offers a partial clue, but in the end you can't decide without running the same task through the models yourself.
So this time, I ran the exact same Japanese question through 7 major models running on a single NVIDIA DGX Spark. That includes NVIDIA's Nemotron 3 Super 120B-A12B, for which I actually deployed the large quantized version.
What I compared wasn't just tokens/s.






