Mark Zuckerberg launched Muse Code in beta on Wednesday, Meta’s first synthetic intelligence (AI) coding agent. Anthropic’s Claude Opus 5 beats it in all 4 comparisons Meta revealed at launch.
Meta launched these charts anyway. The corporate is promoting a less expensive instrument relatively than a greater one. Impartial check knowledge suggests the hole is wider than Meta confirmed.
Comply with us on X to get the newest information because it occurs
Muse Spark 1.2 is the mannequin inside Muse Code. It scored 82.9% on Terminal-Bench 2.1.
Claude Opus 5 scored 86.7% on the identical check. Terminal-Bench comes from the Laude Institute and Stanford researchers. It units 89 actual jobs spanning system restore, knowledge work, and safety.
Second place is respectable. Muse Code beat OpenAI’s Codex at 81.8% and Grok Construct at 81.6%.
The subsequent chart was harsher. DeepSWE 1.1 units 113 coding duties with web entry switched off throughout grading. Muse Spark 1.2 dropped to 3rd at 59.3%.
Meta then revealed a check it constructed itself, drawn from 440 actual pull requests by its personal engineers. Muse Spark 1.2 scored 70.6% there, roughly 9 factors behind Opus 5.
That rating sits solely 2.3 factors above Muse Spark 1.1, the mannequin Meta shipped in July.
Meta measured itself in opposition to GPT-5.6 Terra. OpenAI sells a stronger mannequin known as Sol, and Meta left it out of all three coding charts.
Sol tops the impartial Terminal-Bench 2.1 leaderboard at 89.5%. Opus 5 follows at 89.1%.
Each figures beat the 86.7% Meta reported for Opus 5. Meta picked a weaker setting of its strongest rival and nonetheless completed behind it.
In opposition to Sol, the true chief, Muse Spark 1.2 trails by 6.6 factors relatively than 3.8.
Meta did embrace Sol in a single place. On a graphics processing unit (GPU) kernel activity working previous 1,000 instrument calls, Sol improved on the baseline by 71.2%. Muse Spark 1.2 managed 68.7% and positioned fourth of six.
One caveat cuts the opposite means. Muse Spark 1.2 doesn’t seem on that public leaderboard but, the place solely 26 of 183 tracked fashions have been examined. Its 82.9% stays a Meta determine.
“Muse Spark 1.2 is our subsequent step as we push towards frontier, with bigger, extra succesful fashions on the best way,” Zuckerberg mentioned in a submit.
Zuckerberg Might Quickly Host the Mannequin Beating His Personal
Meta is reportedly in talks to lease compute to Anthropic. The deal may attain $10 billion over two years. Meta knowledge facilities would then assist run the Claude fashions Muse Code was constructed to unseat.
The management behind Muse Code was costly. Zuckerberg paid $14.3 billion in June 2025 for Scale AI and its founder Alexandr Wang, who now heads Meta Superintelligence Labs.
Worth is the lever Wang has left. Charges match the July launch of Meta’s first paid API at $1.25 per million enter tokens and $4.25 per million output tokens.
A contributor tier prices greater than 10 instances much less. Builders qualify by letting Meta practice on their work. Wang declined to present adoption numbers for the Muse Spark line.
Meta’s accounts clarify the low cost. Income climbed 28% to $60.8 billion final quarter, but working revenue fell 8% to $18.8 billion.
Working margin slid to 31% from 43% a 12 months earlier. Meta spent $31.08 billion on capital tasks within the quarter alone, and guides to as a lot as $145 billion for the 12 months.
Muse Code does supply engineering Claude Code lacks. Background brokers maintain context throughout a session. Sub-agents work in remoted copies of a repository.
Meta has constructed a stable second-best coder and priced it like a finances possibility. The beta will present whether or not builders commerce a number of factors of accuracy for a invoice roughly a tenth the scale.
The submit Zuckerberg’s Muse Code Loses to Anthropic on Meta’s Personal Benchmark Charts appeared first on BeInCrypto.
Supply hyperlink