The latest large language models have high false-positive rates and fail to take into account the context of scans, leading to more work for AppSec professionals.

July 21, 2026

Current methods of prioritizing vulnerabilities are falling flat, with too many false positives, poor prioritization, and a failure to take into account reachability.

So far, large language models (LLMs) have not really helped.

In tests of more than a dozen application-scanning tools, more than 60% of flagged vulnerabilities continue to be false positives, are in unreachable code, or are low severity, says Arshan Dabirsiaghi, chief technology officer and co-founder at Pixee, an AI-powered application-security (AppSec) startup. In a presentation at Black Hat USA in August, Dabirsiaghi plans to detail results from those tests and show that the lack of context in stock models means that AI models are not the solution.