Snyk VulnBench JS 1.0 is a repeatability benchmark that ran 300 identical security-review sessions across six LLM configurations to measure whether agentic AI finds the same JavaScript bugs on the second pass as on the first. Snyk published the full results on June 29.
The headline result is not that one scanner outperforms another. It is that LLM findings split sharply by type: when a model matched a reference finding from Snyk Code's deterministic SAST output, it repeated that finding reliably across all five runs of the same task in 85% of cases. When a model reported something outside that reference set, the results were far less stable. Nearly half of those extra reports (80 of 161 unique unmatched signatures) appeared in only one of five identical runs.
Snyk VulnBench JS covered 10 small Express-based JavaScript fixtures with 44 Snyk Code reference findings. Each of the six configurations ran every fixture five times. The scoring method was lenient: a model finding received credit if it matched the same vulnerability type as a reference finding, regardless of file, line, or source-to-sink path. Even under those terms, the best LLM configuration reached 75.4% Snyk-reference F1, leaving a 24.6-point gap against deterministic SAST.
Cost did not predict coverage, according to the Snyk report. Claude Opus 4.7 Max consumed 1.9 times more tokens and cost 5.7 times more per session than Claude Opus 4.6 Medium, while scoring lower on Snyk-reference F1 (68.8% versus 75.4%). Claude Sonnet 4.6 Medium generated the noisiest output; 61.7% of its extra vulnerability reports appeared in just one of five identical scans.
The complementarity finding may carry the most practical weight for engineering teams. Models and SAST did not fail on the same vulnerability classes. Across all configurations, LLMs consistently caught high-signal exploit shapes such as command injection, SQL injection, SSRF, and hardcoded credentials. They were weaker on resource-limit findings, repeated path traversal flows, and framework information exposure, which are precisely the categories where deterministic data-flow analysis has a structural advantage. In the Snyk VulnBench JS benchmark's largest fixture, even the best-scoring model reached only 40.0% Snyk-reference F1 and missed every path traversal reference finding.
The report also surfaces a case where the model was likely right and the reference set was not: all 25 model runs flagged SQL injection in a fixture where the vulnerable-looking code used a logging stub with no executable SQL sink. A separate fixture produced the reverse, with models consistently reporting an unmatched SQL injection finding that Snyk says is a real product gap it plans to address internally. Snyk acknowledges the circular limitation directly: using its own SAST output as the ground truth makes the F1 metric an agreement score, not an accuracy claim.
Planned follow-up work for Snyk VulnBench JS includes larger application structures, an independent externally reviewable ground truth source, business-logic vulnerability classes, and a combined workflow track evaluating SAST-augmented LLM review alongside model-only and SAST-only runs. The core open question is whether combining the two approaches recovers the coverage gap without multiplying triage burden from one-off noisy reports.












