I see a lot of claims about which model is "best." Best at what? For whom? At what cost?
I got tired of guessing. So I ran my own comparison.
The setup
I took 500 real queries from my production logs – a mix of:
Code generation (120 queries)
I see a lot of claims about which model is "best." Best at what? For whom? At what cost? I got tired...
I see a lot of claims about which model is "best." Best at what? For whom? At what cost?
I got tired of guessing. So I ran my own comparison.
The setup
I took 500 real queries from my production logs – a mix of:
Code generation (120 queries)

The Problem With Choosing a Local Model Everyone has an opinion on which local LLM is...

The more models I test, the less comfortable I am answering the question: Which LLM should I use? A...

My app generates personalized readings for BaZi — Chinese "Four Pillars" birth charts. Every reading...

A free-model NL-to-SQL bench scored 17/20, then 6/20 ninety seconds later. The model didn't change — the providers got tired. How…

Originally published on my Substack. I'm a Microsoft MVP based in Japan, writing in English about the...

A while back I got annoyed at a specific genre of blog post: "we asked ChatGPT what the best CRM is...