TUE, JULY 28, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ Large Language Models

Grok 4.5 Wins Vercel DeepsecBench Price-Performance Crown — $5.60 Per Run vs GPT-5.6 Sol at $56

Vercel's DeepsecBench on real cybersecurity vulnerabilities: GPT-5.6 Sol leads accuracy at 35.58 score ($56/run). Grok 4.5 scored 15.58-16.54 at $5.60-$11/run — Vercel CEO Guillermo Rauch named it the price-performance winner. Kimi K3 scores at a comparable level. Caveat: Artificial Analysis found Grok 4.5 hallucination rate doubled (25% to 54%) vs Grok 4.3 — validate before production security use.

By AIToolsRecap July 28, 2026 6 min read 18 views
Home Articles Large Language Models Grok Grok 4.5 Wins Vercel DeepsecBench Price-Perform...

VERCEL DEEPSECBENCH RESULTS — JULY 2026

Benchmark: Vercel DeepsecBench — real cybersecurity flaws in open-source code
Raw accuracy leader: GPT-5.6 Sol — 35.58 score at $56 per run
Grok 4.5: 15.58 to 16.54 score — $5.60 to $11 per run
Kimi K3: Comparable score to Grok 4.5 at similar level
Price-performance winner (per Vercel CEO Guillermo Rauch): Grok 4.5
Cost comparison: GPT-5.6 Sol costs 5-10x more per run for roughly 2x the raw score
Available on Vercel: Yes — Grok 4.5 is one of the third-party gateways listed at launch
Key caveat: Grok 4.5 hallucination rate doubled vs Grok 4.3 (25% to 54%) per Artificial Analysis

What DeepsecBench Actually Tests

DeepsecBench is Vercel's internal benchmark for evaluating AI models on real vulnerability detection in open-source code — not synthetic challenges or CTF problems. The benchmark runs models against actual security flaws found in real projects, measuring whether the model can identify the vulnerability and explain it correctly. Vercel CEO Guillermo Rauch shared the results on X and called Grok 4.5 the top price-performance pick, noting that GPT-5.6 Sol's 35.58 score at $56 per run — roughly 10x the cost of Grok 4.5's $5.60 minimum — does not justify the accuracy delta for teams with high-volume security scanning workloads.

According to eesel AI's Grok 4.5 review, Grok 4.5 is available on Vercel as one of its listed third-party model gateways at launch, alongside OpenRouter, Cloudflare, Snowflake, and Databricks Mosaic. For teams already building on Vercel's infrastructure, using Grok 4.5 for security scanning is a zero-additional-setup change — the model is already accessible through the same gateway they use for other AI features.

The Price-Performance Calculation

Model DeepsecBench score Cost per run Score per $10 spent
GPT-5.6 Sol 35.58 (highest) $56 6.4
Grok 4.5 15.58 – 16.54 $5.60 – $11 ~15 (best)
Kimi K3 ~15-16 (comparable) ~$10-15 (API) ~10

Score per $10 spent calculated from reported benchmark scores and per-run costs. Kimi K3 cost estimate based on $3/M API rate at typical security scan token volume. Vercel CEO named Grok 4.5 as the price-performance winner based on the full results table.

The Grok 4.5 Hallucination Caveat — Critical for Security Use Cases

For cybersecurity scanning specifically, hallucination rate matters more than in most other AI tasks. A model that confidently identifies a non-existent vulnerability wastes engineering time on false positives; a model that misses a real vulnerability creates security risk. According to AIToolsReview UK's independent assessment, Artificial Analysis found that Grok 4.5's hallucination rate more than doubled versus Grok 4.3 — from 25% to 54% — alongside a genuine accuracy gain on other benchmarks. This finding was not published in xAI's own model documentation.

The 54% hallucination rate in general tasks does not directly translate to a 54% false-positive rate on DeepsecBench — the benchmark is domain-specific and the model's coding training may produce better precision on structured vulnerability detection than on open-ended factual recall. But it is a finding that any team considering Grok 4.5 for production security scanning should evaluate on their own dataset before deploying at scale. The price-performance case is real. The hallucination risk requires validation.

What This Means for AI-Assisted Security Scanning

The cost case is real for high-volume scanning: If your team runs security scans across hundreds of repositories daily, the difference between $56 and $5.60 per run is not marginal — it is the difference between AI security scanning being economically viable or not. Grok 4.5 at 10x lower cost per run makes AI-assisted vulnerability detection accessible for teams that cannot justify GPT-5.6 Sol pricing at that volume.

Validate before production deployment: The 54% hallucination rate finding from Artificial Analysis is relevant. Run Grok 4.5 against a sample of your known vulnerabilities and known-clean code before deploying as a primary scanner. The DeepsecBench result is encouraging but single-benchmark evidence.

Grok 4.5 is available on Vercel now: No additional API setup required for Vercel users. The model is accessible through the same Vercel gateway already used for other AI features. According to o-mega.ai's launch coverage, Grok 4.5 is live on Vercel alongside OpenRouter, Cloudflare, Snowflake, and Databricks at launch.

Sources: Vercel DeepsecBench results via Guillermo Rauch X · eesel AI — Grok 4.5 review · AIToolsReview UK — hallucination finding · o-mega.ai — availability · Related: Grok 4.5 full review · Claude Opus 5 vs Grok 4.5

Tags
GrokAI NewsGenerative AICoding AI2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →