THU, SEPTEMBER 24, 2026
Independent · In‑Depth · Practitioner‑Tested
Large Language Models

Claude Opus 5.5 vs GPT-6 Astra on Computer Use: The Numbers Do Not Compare

One model scores 81.8% and 48.7% on the same benchmark depending on who measured it.

🕐 7 min read 👁 13 views 📅 Sep 24, 2026

The Short Version

You cannot currently choose between these two models on their computer-use scores, because no two published figures were produced the same way. This comparison is about why, and what to do instead.

The Same Model, Three Numbers

Model Score Who measured it, and how
Claude Opus 5.5 81.8% Anthropic, OSWorld 2.0 with partial credit, own harness
Claude Opus 5.5 48.7% BenchLM public leaderboard, 23 Sept - 8th place
GPT-6 Astra 72.6% OpenAI, offline task subset - and 1st on the same leaderboard
Claude Opus 5 74.0% Anthropic's measurement
Claude Opus 5 70.2% OpenAI's measurement of the same model

Read That Table Again

Three things in it should stop you.

One model, a 33-point spread. Opus 5.5 is 81.8% or 48.7% depending on who ran it. Neither party is lying - they are running different task sets with different scoring rules, and partial credit against strict pass/fail will move a number that far on a long-horizon benchmark.

The vendors disagree about a model neither of them is selling any more. Anthropic scores Opus 5 at 74.0%, OpenAI scores it at 70.2%. That 3.8-point gap, on a settled model, is larger than the gap between several models people are currently choosing between.

The rank inversion. On Anthropic's numbers Opus 5.5 beats Astra by nine points. On the public leaderboard Astra is first and Opus 5.5 is eighth. Same two models, opposite answer.

What OSWorld 2.0 Actually Tests

108 long-horizon computer-use workflows requiring state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction and verification. Long-horizon is the operative word: the longer the task, the more scoring methodology matters, because partial credit on a 30-step task is a judgement call about what counts as progress.

That is why this benchmark diverges more than, say, a coding benchmark where the tests pass or they do not.

The Benchmarks That Do Compare

Not everything this week is unusable. Where both vendors published on the same task set, the figures are worth reading:

  • Terminal-Bench 4.0: Opus 5.5 66.4%, Astra 57.9%
  • GDPval-AA v2.1: Opus 5.5 1846 Elo, Astra 1542
  • Humanity's Last Exam, with tools: Opus 5.5 67.7%, Astra 57.2%
  • FrontierCode v1.1: Opus 5.5 54.4%, Astra 53.3%
  • AutomationBench: Opus 5.5 40.0%, Astra 41.4%
  • Terminal-Bench-Science 0.1: Opus 5.5 58.7%, Astra 64.6%

Opus 5.5 takes four, Astra takes two, and two of the four are inside a point and a half. The honest summary is that they are close on everything measurable and Terminal-Bench is the one real gap.

The One Clean Signal This Week

DrivingBench ran four models through one physical course on one harness with pass or fail scoring. Astra finished, Claude Fable 5.1 got 45%, Grok 4.6 got 11%.

It is a much smaller result than OSWorld and a much more trustworthy one, because a car either completes a course or it does not, and one party ran all four models. When vendor benchmarks diverge this far, small clean measurements beat large contested ones.

Decision Framework

  • Choosing a model for computer use - do not choose on OSWorld. Build a 20-task harness from your own workflows and run both. It is a day of work and it is the only number that will be true for you.
  • Terminal and shell automation - Opus 5.5, on the 8.5-point Terminal-Bench gap, which both vendors measured on the same set.
  • Browser automation specifically - genuinely unresolved on published figures. Test it.
  • Cost is the constraint - Opus 5.5 at $4/$20 against Astra at $10/$50. The capability difference across comparable benchmarks does not justify 2.5x.
  • Quoting a benchmark to a stakeholder - name the source and the scoring method in the same sentence, or you will be quoting noise with a decimal point on it.

Verdict

On comparable benchmarks these two models are close, Opus 5.5 leads terminal work, and Opus 5.5 costs 60% less. On computer use specifically, there is no published figure either of us should be acting on.

The wider point is worth more than the model choice: when a single model can be scored 81.8% or 48.7% on the same benchmark name, benchmark citation without methodology is not evidence. Run your own twenty tasks.

⚖ Our Verdict

On comparable benchmarks the two are close, Opus 5.5 leads Terminal-Bench by 8.5 points, and it costs 60% less. On computer use there is no usable published figure - Opus 5.5 is scored 81.8% by Anthropic and 48.7% on the public leaderboard, and the vendors disagree by 3.8 points about Opus 5. Build your own 20-task harness.