Evaluating LLMs on standardized leaderboards (like MMLU or HumanEval) is helpful,
but it rarely tells you how a model performs on real-world edge cases.
In this benchmark, I tested three models on a specific dev-sec scenario:
Detecting hidden reentrancy and integer overflow vulnerabilities in a complex smart contract.
Models Tested:








