I used to pick coding models the same way people pick sports cars: choose the most powerful one and pretend the fuel bill is somebody else's problem.

That worked when I was asking one question at a time. Then I started using agents for real engineering work.

An engineering agent does not answer once and disappear. It reads the repository, searches for related code, opens the wrong file, finds the right file, proposes a change, runs a test, breaks something, reads the error, fixes the change, and runs the test again. Sometimes it also writes a surprisingly thoughtful essay about the three lines it just modified.

By the time one task is finished, the model may have been called a dozen times. Suddenly, model pricing is not a footnote. It is part of the architecture.

That is how GPT-5.6 Luna with high reasoning effort became my default.