Last Updated: September 3, 2026. Covers the Muse Spark 1.3 release of September 2, 2026, from Meta Superintelligence Labs.
Meta released Muse Spark 1.3 on September 2, 2026, and it is the company's most consequential update of the year for anyone building AI agents. According to Meta AI Research, the model delivers Meta's biggest jump yet on coding and agentic tasks: 75.4% on DeepSWE 1.1 for end-to-end agentic software engineering, 88.8% on Terminal-Bench 2.1, 59.4% on SWEAtlas CodeBase QnA, and 98.5% on long-context retrieval, all inside a 1 million token context window. Meta's engineers measured it finishing coding work with roughly 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2. The honest read: Muse Spark 1.3 now wins coding and long-context work outright, while Claude Opus 5 still leads the general knowledge-work and autonomy rows in Meta's own table. The bigger signal is strategic. Meta's first closed, directly monetised frontier line is shipping on a four-week cadence and pressuring every rival on price.
Key takeaways
- Released September 2, 2026: same-day availability in Muse Code and the Meta Model API, with paid developer access opening September 3.
- Benchmark wins: 75.4% DeepSWE 1.1, 88.8% Terminal-Bench 2.1 (tied with GPT-5.6 Sol), 59.4% SWEAtlas CodeBase QnA, 98.5% MRCR at 256k to 512k.
- Efficiency: about 20% fewer tool calls and 25% fewer tokens versus Muse Spark 1.2, per comparisons by Meta engineers.
- Agentic behaviour: asks clarifying questions, confirms before consequential actions, and is better calibrated on irreversible steps.
- Where it trails: JobBench, OSWorld 2.0, AutomationBench and GDPval-AA v2 each sit behind Claude Opus 5 on Meta's own comparison table.
- Strategic shift: proprietary weights, a paid API, and an open weights release promised on the roadmap with no date attached.
What Is Meta Muse Spark 1.3?
Muse Spark 1.3 is a proprietary multimodal reasoning model from Meta Superintelligence Labs (MSL), the unit assembled around Alexandr Wang after Meta's $14.3 billion investment in Scale AI. It carries a 1 million token context window and is built specifically for long-running agentic, multi-agent and coding workflows: sustaining work across extended tasks, keeping track of what it has learned, and juggling multiple workflows inside a single long thread. It shipped September 2, 2026 into Muse Code, Meta's coding agent for the terminal and CI, and into the Meta Model API. The weights are closed; paid API access opened the following day.
The Muse line has moved fast. The original Muse Spark launched on April 8, 2026, internally codenamed Avocado, as the first model out of MSL and the product of a nine-month ground-up rebuild of Meta's entire AI stack: new architecture, new data pipelines, and infrastructure anchored by the Hyperion data centre. It was a hard break from the Llama family. According to NeuralCoreTech's launch analysis, the April model scored 52 on the Artificial Analysis Intelligence Index v4.0 against 57 for GPT-5.4 and Gemini 3.1 Pro, led every rival on HealthBench Hard at 42.8, and completed the full index evaluation using 58 million output tokens, roughly a third of Claude Opus 4.6's 157 million, thanks to a technique Meta calls thought compression. Muse Spark 1.1 followed, then version 1.2 on August 5, 2026, co-trained with the newly launched Muse Code. A second Meta model, Muse Glimmer, arrived on August 10. Four weeks after 1.2 came 1.3, announced not with a launch post but by Mark Zuckerberg on X, pitched, according to AI Release Tracker, as frontier performance "almost too cheap to meter."
One framing matters for the rest of this analysis. "Muse Spark 1.3 advances our work toward personal superintelligence," Meta AI Research wrote in its announcement. That is not throwaway language. Every design choice in this release, from token efficiency to agentic caution, points at a model line intended to run continuously for billions of people rather than to top a leaderboard once.
What Do the Muse Spark 1.3 Benchmarks Actually Show?
Meta published 11 benchmark scores at release, spanning coding, long context, knowledge work and agentic tool use. Muse Spark 1.3 posts the best tracked figures on agentic coding (DeepSWE 1.1: 75.4%), codebase comprehension (SWEAtlas CodeBase QnA: 59.4%) and long-context retrieval (MRCR: 98.5% at 256k to 512k, 98.1% at 512k to 1M). It ties GPT-5.6 Sol on Terminal-Bench 2.1 at 88.8%. The knowledge-work rows tell the other half of the story: JobBench, OSWorld 2.0, AutomationBench and GDPval-AA v2 each sit just behind Claude Opus 5, per Meta's own comparison table.
| Benchmark | Muse Spark 1.3 | What it measures | The read |
|---|---|---|---|
| DeepSWE 1.1 | 75.4% | Deep, agentic software engineering end to end | Best of tracked models |
| SWEAtlas CodeBase QnA | 59.4% | Understanding an unfamiliar codebase | Ahead of every rival on Meta's table |
| Terminal-Bench 2.1 | 88.8% | Command-line and technical setup work | Tied with GPT-5.6 Sol |
| MRCR (256k to 512k) | 98.5% | Finding details buried in very long inputs | Up from 66.3% on Muse Spark 1.2 |
| MRCR (512k to 1M) | 98.1% | Same test on even longer spans | Up from 55.5% on Muse Spark 1.2 |
| JobBench | 64.9% | Multi-step workplace tasks with real tools | Just behind Claude Opus 5 |
| OSWorld 2.0 | 66.9% | Operating a computer: clicking, typing, real apps | Just behind Claude Opus 5 |
| AutomationBench | 49.4% | End-to-end business workflow automation | Behind Claude Opus 5 |
| GDPval-AA v2 | 1754 | Economically valuable knowledge work (Elo) | Behind Claude Opus 5 |
| DeepSearchQA | 89.4% | Multi-source web research and synthesis | Trails both rivals |
| Agentic IF Index | 57.8% | Instruction fidelity across long agent runs | Trails both rivals |
Three things in that table deserve attention. First, the long-context jump is enormous: 66.3% to 98.5% and 55.5% to 98.1% between 1.2 and 1.3, while Claude Opus 5 reported no figure at either span on Meta's table. A model that reliably retrieves detail from several books' worth of context changes what you can build: whole contract sets, full codebases and complete patient or client histories in one thread. Second, the coding story has flipped. In April, the original Muse Spark posted 52.4% on SWE-Bench Verified while Claude Opus 4.6 led at 80.8%, per NeuralCoreTech's data. Five months later Meta leads the agentic coding rows. Third, read the footnotes before celebrating: according to AI Release Tracker, the comparison table ran Muse Spark 1.3 at its max reasoning setting against Muse Spark 1.2 at xhigh, and the rival figures are Meta's own runs of GPT-5.6 Sol and Claude Opus 5 rather than those labs' published numbers, with several differing from what OpenAI and Anthropic report themselves. Bloomberg's launch coverage framed the release as Meta's most powerful model yet, with its chief AI officer saying capabilities are edging closer to top competitors. Closer, not ahead. That framing matches the table.
How Does Muse Spark 1.3 Change AI Agents?
Muse Spark 1.3 is trained explicitly for long-horizon agency: sustaining multi-step work in a single thread, generating its own context from messy and conflicting sources, correcting gaps in its own plan, and tracking what it has learned toward a final deliverable. It asks clarifying questions when prompts are ambiguous, requests help when stuck, and confirms before consequential actions. Combined with stronger resistance to prompt injections and better calibration on irreversible steps, this is the most agent-shaped release Meta has shipped: it attacks the two things that actually block agent deployments in real businesses, reliability over long runs and trust in consequential moments.
The behaviour changes matter more than the scores. Meta describes a model that maps incoming prompts to the correct task inside messy, single-threaded conversations, whether the user is steering past requests or interrupting them. Anyone who has run an agent through a real workday knows the failure mode this fixes: the agent conflates two tasks, silently drops a constraint, or plows ahead on stale instructions. Meta also trained better self-knowledge: what the model can and cannot do, what it knows, and when it has hit a hurdle "instead of hallucinating outcomes," in the company's words. An agent that reports being stuck is an agent you can actually hand work to.
The safety work is the quiet headline for agent builders. According to Meta AI Research, Muse Spark 1.3 shows "stronger adversarial robustness, with improved resistance to adversarial inputs and prompt injections," plus better calibration on what constitutes an irreversible action. Prompt injection is the number one attack vector for tool-using agents: a poisoned web page or email that tells your agent to exfiltrate data or fire off messages. A frontier lab training defences against that directly into the model, and gating its max reasoning mode behind additional safety testing, signals that agent security is becoming a first-class model feature rather than an add-on wrapper. It does not replace your own approval gates, audit logs and dedup checks. It lowers the baseline risk of every agent you deploy on it.
Then there is economics. Meta's engineers measured Muse Spark 1.3 using roughly 20% fewer tool calls and 25% fewer tokens than 1.2 on the same engineering workflows. At Flowtivity we run always-on agents across research, pipeline updates and client reporting, and we re-quote our model routing roughly monthly because metered API cost is a real line item. A 25% token reduction at frontier quality compounds fast in that world: it is the difference between an agent that watches costs and an agent you let run. Combine that with the line's inherited efficiency, the original Muse Spark completed a full frontier evaluation on 58 million output tokens where Claude Opus 4.6 used 157 million, and the multi-agent Contemplating mode that beat GPT-5.4 Pro and Gemini Deep Think on Humanity's Last Exam at 50.2%, and you have a stack built to make always-on, multi-agent work cheap. That is precisely what agent businesses need.
What Does Muse Spark Mean for the Wider AI Industry?
Muse Spark matters to the industry less as a leaderboard entry and more as a business-model pivot. Meta spent years giving away frontier-adjacent models through Llama, and that approach is quietly ending: Muse Spark is closed, proprietary, and directly monetised through the Meta Model API, whose paid tier opened September 3, 2026. The release also landed in a week where, according to Axios, Meta is "trying to keep pace with Anthropic, OpenAI and Google who all made major model announcements." Four frontier-quality Muse Spark releases in five months, plus Muse Glimmer, is a cadence the industry has only seen from OpenAI before.
The pricing signal is the sharpest part. Zuckerberg pitching frontier performance as "almost too cheap to meter" is not an idle boast when your architecture is demonstrably token-efficient; it is a margin attack. Thought compression means Meta's inference cost per query is structurally lower than rivals that burn two to three times the output tokens for equivalent work, and Meta can pass that through to API pricing while still funding free consumer access across an app base of more than 3 billion daily users. For the industry, expect margin compression in the API layer, another leg down in per-token prices, and renewed pressure on smaller model providers. For agent builders, falling token prices are pure tailwind: the cost of a competent always-on agent keeps dropping faster than capability rises, which is the actual mechanism pulling agents into every business process.
The open question is open weights. Meta explicitly lists "the Muse Spark open weights release" on its roadmap alongside bigger models, with no dates. That is a careful straddle: keep the monetised frontier closed, promise openness later, and blunt the criticism that Meta abandoned the open ecosystem it helped build. Watch this one closely. If a recent Muse Spark ships as open weights, it resets the price floor for self-hosted agents globally and hands every integrator and consultancy a free frontier-class option. If it keeps sliding, the open-source frontier stalls at Llama's last generation and Meta's pivot becomes the template every lab follows. Also watch for independent verification: the competitive figures in Meta's table are its own runs, and third-party trackers like Artificial Analysis had the original Muse Spark five points behind GPT-5.4 and Gemini 3.1 Pro in April. Independent numbers on 1.3 will shape the real narrative.
What Should Businesses Building AI Agents Do Now?
If you run agents in production, three moves are worth making this month. Route per task rather than picking one model: Muse Spark 1.3 is now the strong candidate for coding agents, long-document analysis and terminal workflows, while Claude Opus 5 remains the pick for general autonomy and computer use, and cost-heavy bulk work still belongs on the cheapest per-token tier. Trial Muse Code on a real repo: it installs on macOS and Linux and the 1.3 model is live in it today. And re-quote your stack: between the token efficiency gains and Meta's price pressure, agent running costs are falling fast enough that last quarter's budget assumptions are stale.
For Australian growing businesses, the practical impact lands in the workflows we build every week: quoting, inbox triage, follow-up sequences, report drafting, pipeline hygiene. Muse Spark 1.3's combination of a 1M token window and strong instruction retention means whole-of-history work becomes reliable: an agent that has read every email, quote and job file for a client, and still follows the brief 60 tool calls later. The improved calibration on irreversible actions matters too. We still wrap every consequential step, sends, payments, deletions, in human approval gates regardless of model, and you should too. But a model that confirms before acting and admits when stuck makes those guardrails cheaper to run and less annoying to staff.
One planning note from our side of the fence: the leaderboard now moves roughly every four weeks. Model choice has become a routing decision you revisit monthly, not an annual platform bet. Budget for evaluation time, keep your agent harness model-agnostic, and treat every benchmark table, including Meta's, as a vendor's own marketing until independent numbers land.
Muse Spark 1.3 FAQ: Quick Answers
When was Muse Spark 1.3 released?
September 2, 2026, four weeks after Muse Spark 1.2, with same-day availability in Muse Code and the Meta Model API and paid API access opening September 3.
Is Muse Spark 1.3 open source?
No. The weights are unpublished and the model is available only through Meta's API and products. Meta lists an open weights release for the Muse Spark line on its roadmap, with no date committed.
What is the context window?
1 million tokens, with 98.5% retrieval accuracy at the 256k to 512k span and 98.1% at 512k to 1M on MRCR.
Is it better than Claude Opus 5?
On coding, codebase comprehension, terminal work and long context, yes, per Meta's own table. On job execution, computer use, business workflow automation and knowledge work, Claude Opus 5 still leads. Route per task.
How much does it cost?
Paid access opened September 3 through the Meta Model API. No premium framing: Zuckerberg pitched it as frontier performance almost too cheap to meter, and Meta's token-efficiency advantage suggests sustained price pressure industry wide.
The Bottom Line on Muse Spark 1.3
Muse Spark 1.3 is not the best model at everything, and Meta's own table says so. It is the model that makes agentic coding and long-context agent work cheap, efficient and safer to run, shipped on a four-week cadence by a company that has decided to monetise the frontier directly. For agent builders, falling token costs, trained-in injection resistance and calibrated caution on irreversible actions are the three changes that reach your P&L and your risk register this quarter. For the industry, the release confirms the new shape of the race: four frontier labs shipping within days of each other, prices trending toward the floor, and open weights as the unresolved question that decides who gets to build for free. Our take after five months of Muse Spark: route per task, verify independently, and enjoy the discount.