The Short Version
You cannot currently choose between these two models on their computer-use scores, because no two published figures were produced the same way. This comparison is about why, and what to do instead.
The Same Model, Three Numbers
| Model |
Score |
Who measured it, and how |
| Claude Opus 5.5 |
81.8% |
Anthropic, OSWorld 2.0 with partial credit, own harness |
| Claude Opus 5.5 |
48.7% |
BenchLM public leaderboard, 23 Sept - 8th place |
| GPT-6 Astra |
72.6% |
OpenAI, offline task subset - and 1st on the same leaderboard |
| Claude Opus 5 |
74.0% |
Anthropic's measurement |
| Claude Opus 5 |
70.2% |
OpenAI's measurement of the same model |
Read That Table Again
Three things in it should stop you.
One model, a 33-point spread. Opus 5.5 is 81.8% or 48.7% depending on who ran it. Neither party is lying - they are running different task sets with different scoring rules, and partial credit against strict pass/fail will move a number that far on a long-horizon benchmark.
The vendors disagree about a model neither of them is selling any more. Anthropic scores Opus 5 at 74.0%, OpenAI scores it at 70.2%. That 3.8-point gap, on a settled model, is larger than the gap between several models people are currently choosing between.
The rank inversion. On Anthropic's numbers Opus 5.5 beats Astra by nine points. On the public leaderboard Astra is first and Opus 5.5 is eighth. Same two models, opposite answer.
What OSWorld 2.0 Actually Tests
108 long-horizon computer-use workflows requiring state tracking, cross-source reasoning, visual-spatial precision, dynamic interaction and verification. Long-horizon is the operative word: the longer the task, the more scoring methodology matters, because partial credit on a 30-step task is a judgement call about what counts as progress.
That is why this benchmark diverges more than, say, a coding benchmark where the tests pass or they do not.
The Benchmarks That Do Compare
Not everything this week is unusable. Where both vendors published on the same task set, the figures are worth reading:
- Terminal-Bench 4.0: Opus 5.5 66.4%, Astra 57.9%
- GDPval-AA v2.1: Opus 5.5 1846 Elo, Astra 1542
- Humanity's Last Exam, with tools: Opus 5.5 67.7%, Astra 57.2%
- FrontierCode v1.1: Opus 5.5 54.4%, Astra 53.3%
- AutomationBench: Opus 5.5 40.0%, Astra 41.4%
- Terminal-Bench-Science 0.1: Opus 5.5 58.7%, Astra 64.6%
Opus 5.5 takes four, Astra takes two, and two of the four are inside a point and a half. The honest summary is that they are close on everything measurable and Terminal-Bench is the one real gap.
The One Clean Signal This Week
DrivingBench ran four models through one physical course on one harness with pass or fail scoring. Astra finished, Claude Fable 5.1 got 45%, Grok 4.6 got 11%.
It is a much smaller result than OSWorld and a much more trustworthy one, because a car either completes a course or it does not, and one party ran all four models. When vendor benchmarks diverge this far, small clean measurements beat large contested ones.
Decision Framework
- Choosing a model for computer use - do not choose on OSWorld. Build a 20-task harness from your own workflows and run both. It is a day of work and it is the only number that will be true for you.
- Terminal and shell automation - Opus 5.5, on the 8.5-point Terminal-Bench gap, which both vendors measured on the same set.
- Browser automation specifically - genuinely unresolved on published figures. Test it.
- Cost is the constraint - Opus 5.5 at $4/$20 against Astra at $10/$50. The capability difference across comparable benchmarks does not justify 2.5x.
- Quoting a benchmark to a stakeholder - name the source and the scoring method in the same sentence, or you will be quoting noise with a decimal point on it.
Verdict
On comparable benchmarks these two models are close, Opus 5.5 leads terminal work, and Opus 5.5 costs 60% less. On computer use specifically, there is no published figure either of us should be acting on.
The wider point is worth more than the model choice: when a single model can be scored 81.8% or 48.7% on the same benchmark name, benchmark citation without methodology is not evidence. Run your own twenty tasks.