The Short Version
Everyone compares coding agents on benchmarks. Benchmarks tell you how a tool behaves on a good day.
In September, seven of them received the same vulnerability report from the same researcher at roughly the same time. What each vendor did next is public, verifiable, and tells you something a SWE-bench score cannot: how this company behaves when it has a problem and you are exposed to it.
The Vulnerability, Briefly
Git's core.fsmonitor setting names a command Git runs during file status checks. It lives in the repository's own .git/config. Point a coding agent at a repository someone sent you and it runs that command - no sandbox, no approval prompt. core.hooksPath and attr.tree behave the same way.
Manifold Security published it as GitSpawn on 1 September 2026. Full technical detail is in our write-up of the disclosure.
The Response Table
| Agent |
What the vendor did |
Grade |
| goose |
Patched in 1.44.0, CVE-2026-72718 published with a CVSS score |
Good |
| Codex |
CLI patched in 0.131.0, Desktop patched, three CVEs self-disclosed |
Good |
| Cursor |
Patched following an earlier disclosure |
Good |
| Claude Code |
fsmonitor path patched in 2.1.196; second path still live in 2.1.258 |
Partial |
| Qwen Code |
Report accepted 7 July. No release. |
Poor |
| Hermes Agent |
Six contact attempts, advisory untriaged |
Poor |
| Grok Build |
Closed as informative; addressed on social media only |
Poor |
Reading the Table Fairly
What "Good" actually means here
goose and Codex both shipped a fix and filed a CVE. The CVE is the part that matters - it means the fix is discoverable by anyone running a vulnerability scanner, rather than a quiet version bump you would only notice if you read changelogs.
The Claude Code case is the instructive one
Anthropic patched quickly - 2.1.196 closed the reported path. Then a second path through "claude ultrareview" was found and was still live in 2.1.258.
This is not negligence; it is the normal shape of a vulnerability class. The first report names one entry point, the researcher keeps looking, and more turn up. What it means practically is that "we patched it" and "it is fixed" are different claims, and if you run Claude Code you should apply the config mitigation rather than trusting the version number.
Why "closed as informative" is the worst answer
Grok Build closing the report as informative is a judgement that this is expected behaviour rather than a flaw. That is a defensible position to argue - and then four other vendors shipped patches for the same thing, which makes it a lonely one. Addressing it on social media instead of in an advisory means users who do not follow the right account never learn about it.
What This Does Not Tell You
- One disclosure is one data point. A vendor that handled this badly may handle the next one well, and vice versa.
- Nothing here is being exploited. No real-world exploitation is documented and none of these CVEs are in CISA's KEV catalog.
- Response speed is not code quality. goose patching fastest does not make it the better agent for your work.
Treat this as one input among several, not a ranking.
Decision Framework
- You work on client code or handle received repositories - use a patched agent and apply
git config --global core.fsmonitor false anyway. Belt and braces.
- Your organisation has a security review - goose and Codex have published CVEs, which is what a reviewer wants to see. An agent with no advisory trail is harder to get approved.
- You use Claude Code - keep using it, apply the config mitigation, and do not assume the latest version closes it.
- You use Qwen Code, Hermes or Grok Build - apply the mitigation yourself today, because the vendor has not.
- You only ever clone from GitHub yourself - your exposure is low. The attack needs a repository that arrives with its
.git directory intact.
Verdict
On disclosure handling: goose and Codex first, both shipping fixes with published CVEs. Cursor patched. Claude Code responded fast and the job is not finished. Qwen Code, Hermes Agent and Grok Build have had this for weeks and done nothing users can act on.
None of that decides which agent writes better code. It does tell you what happens the next time one of them has a problem while your repository is open in it - and that is worth knowing before you find out the hard way.