What this article covers: A measured run of Anthropic's official security-scanning plugin claude-security (beta). What the tool does, how long it takes, what output it returns, and how far you can trust it — backed by the raw data from the run's own artifacts.
You may know the official plugin exists and still not have tried it, for reasons like these:
The cost is unreadable. It warns you at startup that it "may take a while and use a significant number of tokens," but never says how much
The output quality is unknown. How is this different from existing static analysis, and does an LLM that only reads code produce findings you can act on?
It's unclear what to do afterward. If dozens of findings come back, where do you start?










