SAT, JULY 25, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ Large Language Models

Meta Muse Spark 1.1 Review: $1.25/$4.25/M, 1M Context, and the End of Meta's Open-Only Era

Meta launched Muse Spark 1.1 on July 9, 2026 — same day as GPT-5.6 Sol and Grok 4.5. $1.25/$4.25/M, 1M context with active self-compaction, parallel sub-agents, Computer Use across desktop/browser/mobile. The real story: Meta's first paid developer API — ending its open-weights-only era. Drop-in compatible with OpenAI and Anthropic SDKs. Terminal-Bench compute cap dispute is the main benchmark caveat. US-only API preview.

By AIToolsRecap July 25, 2026 7 min read 23 views
Home Articles Large Language Models Meta Muse Spark 1.1 Review: $1.25/$4.25/M, 1M C...

META MUSE SPARK 1.1 — KEY SPECS (LAUNCHED JULY 9, 2026)

API pricing: $1.25/M input · $4.25/M output
Free credits: $20 on sign-up
Context window: 1M tokens (1,048,576) with active self-compaction
Type: Multimodal reasoning model — text, image, video, PDF
Agentic features: Parallel sub-agents, Computer Use (desktop, browser, mobile), tool-calling
Access: Meta AI app, meta.ai (consumer), Meta Model API (developer preview, US-only at launch)
SDK compatibility: Drop-in for OpenAI and Anthropic SDKs
Benchmarks (SOTA claimed): MCP Atlas, JobBench, Humanity's Last Exam, FinanceBench
Benchmark caveat: Terminal-Bench 2.1 compute cap dispute (see below)
Weights: Closed (first closed Meta frontier model)
FLI Safety Index: D+ (Meta excluded from White House 30-day AI review framework)

The Real Story: Meta's First Paid API

On July 9, 2026, Meta released Muse Spark 1.1 and put it behind its first paid developer API. For a company whose public identity in AI was built on giving Llama weights away, the launch of the Meta Model API is a genuine strategic turn. Meta released Muse Spark 1.1 on July 9, 2026, and put it behind its first paid developer API. The benchmark story matters, but the API story is what actually changes the competitive landscape. Meta has spent years being the "open-weights" company. A hosted, metered, paid API targeting the same enterprise developers as OpenAI and Anthropic is a different business.

Mark Zuckerberg amplified the release personally on X, leaning on coding and agentic performance as the pitch. When the CEO promotes a mid-point version bump himself, it is a tell about where Meta thinks the competitive fight is right now: not chat, but agents that write code and drive software. Muse Spark 1.1 is also notable as the first Meta model that is closed-weight. The open Llama family continues, but Muse Spark operates on Anthropic and OpenAI's model: proprietary weights, hosted inference, API pricing.

Pricing in Context — $1.25/$4.25/M

ModelInput /1MOutput /1MContext
Meta Muse Spark 1.1$1.25$4.251M
Claude Opus 5$5$251M
GPT-5.6 Sol$5$301M
Gemini 3.6 Flash$1.50$7.501M

API pricing is $1.25 per million input tokens and $4.25 per million output tokens, with $20 in free credits. The pricing puts it slightly above Claude Haiku 4.5 and GPT-5.6 Luna — a low-cost frontier tier positioning, not an Opus 4.8-tier price. At $4.25/M output Muse Spark 1.1 is significantly cheaper than Claude Opus 5 ($25/M) or GPT-5.6 Sol ($30/M). The positioning is clear: better than the budget tier, cheaper than the frontier — targeting teams who want agentic capability without frontier-tier pricing.

What It Can Do — Agentic Features

Parallel sub-agents: Muse Spark 1.1 can act as a main agent that creates a plan and delegates work to parallel subagents. This is the architecture that makes multi-step agentic workflows faster — the orchestrator spawns specialist sub-agents for subtasks rather than running everything sequentially.

Active context management: The 1M-token context window is actively compacted by the model itself. Rather than filling the context window until it hits the limit, the model compacts earlier steps while preserving what it will need later — a practical architecture for long-horizon agentic tasks.

Computer Use: Desktop, browser, and mobile surface support. Can interact with GUI applications, fill forms, navigate browsers, and control desktop apps — similar to Claude's Computer Use capability.

OpenAI/Anthropic SDK drop-in: OpenAI and Anthropic SDK compatibility makes an A/B test cheap. Change the base URL and model string — no SDK rewrite needed.

The Terminal-Bench 2.1 Dispute — The Biggest Asterisk

The compute cap dispute over Terminal-Bench 2.1 is the single largest asterisk on the headline benchmark story. Meta ran with 6 CPU cores and 8GB RAM caps where 0 of 89 tasks allow 6 cores and only 8 of 89 allow 8GB. Hacker News flagged this immediately after launch. The practical implication: if Meta's Terminal-Bench run used constrained compute settings that disadvantaged competing models, the headline scores overstate Muse Spark 1.1's relative performance on that benchmark. Muse Spark 1.1 appears strongest as an agent and workflow model, competitive but not dominant as a coding model, mixed on pure long-context retrieval, strong in tool-augmented reasoning.

Honest Limitations

Closed weights: No local deployment, no fine-tuning. For teams that need self-hosted inference or custom fine-tunes, Muse Spark 1.1 is not an option. The open Llama family still exists but is a different product.

US-only preview: The Meta Model API launched as US-only. Global availability timeline not confirmed at launch.

Coding trails Opus 4.8 and GPT-5.5: Muse Spark 1.1 tops Meta's reported tool-use evals but trails Opus 4.8 and GPT-5.5 on coding. The agentic orchestration story is real; the "best coding model" claim requires the Terminal-Bench caveat.

FLI Safety Index D+: Meta scored D+ on the FLI Summer 2026 AI Safety Index — and is not included in the White House voluntary 30-day pre-release AI review framework. The lowest safety governance score among the major Western labs.

Who Should Use Muse Spark 1.1

Agentic workflow builders (US): If you are building multi-step agent workflows and want parallel sub-agent delegation at $4.25/M output — significantly cheaper than Claude Opus 5 or GPT-5.6 Sol — Muse Spark 1.1 is worth evaluating. The SDK drop-in compatibility means the test is cheap.

Teams already using Meta AI / WhatsApp: If your users are in Meta's consumer ecosystem, Muse Spark 1.1 in the Meta AI app gives you the same model your backend uses.

Not recommended for: Production coding requiring verified benchmark accuracy, regulated industries where Meta's D+ FLI safety score or exclusion from the White House framework creates governance risk, and any team outside the US (API preview is US-only).

Sources: Meta official announcement · DataCamp · Neowin · MarkTechPost · ExplainX.ai · Kie.ai · TheAIDude · Metirai · Related: FLI Safety Index — Meta D+ → · White House framework (Meta excluded) → · AI safety comparison →

Tags
AI NewsGenerative AIChatGPT2026AI agents

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →