Evaluating LLMs on standardized leaderboards (like MMLU or HumanEval) is helpful,

but it rarely tells you how a model performs on real-world edge cases.

In this benchmark, I tested three models on a specific dev-sec scenario:

Detecting hidden reentrancy and integer overflow vulnerabilities in a complex smart contract.

Models Tested: