Late on a Tuesday, a C++ tooling team noticed their warning classifier was duplicating classifications. The batch runner sent forty compiler diagnostics to a free model endpoint. The first eight came back fine. The ninth returned the same JSON as the fourth. The tenth timed out. The retry loop sent the ninth again. The endpoint answered with a different label. The night ended with a half-written warning database and a confused engineer.
The service was small. A cron job ran clang-tidy on a legacy codebase. It collected warnings. It sent each warning to a free model endpoint with a prompt asking for a severity label: real, noise, or review. The model was supposed to reduce the queue of warnings a human had to inspect. A health check passed before the batch. The team assumed a passing health check meant the endpoint would behave. It did not.
They configured the harness with two settings. One pointed at MonkeyCode's free model access. The other pointed at the free server option. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The failure shapes were different. The first setup returned stale duplicates under load. The second setup dropped connections after a successful warm-up. Neither problem was solved by adding another retry.






