Card: Meta's best Muse mode is still in testing — Spark 1.3 is live. Max reasoning and open weights are not.

Meta's new Muse model is available today. The reasoning mode at the top of Meta's own comparison table is not.

What shipped. Meta says Muse Spark 1.3 is rolling out in Muse Code and the Meta Model API. The modes that were already available remain available. A new setting called max reasoning will come “shortly” after additional safety testing, but Meta gives no date or detail about what that testing must establish.

The release is more than a benchmark chart. Meta says 1.3 can keep several tasks straight in one long conversation, ask clarifying questions, request help when stuck, and confirm before consequential actions. Those are vendor claims, not yet proof that the behavior holds across real products, but they name the right failure modes for agents doing extended work.

What Meta compared. The evaluation methodology says Meta used max reasoning for its Muse Spark 1.3 comparison and xhigh for Muse Spark 1.2. The full table also includes a Muse Spark 1.3 xhigh column, so it does permit a same-setting comparison with 1.2 xhigh. Meta's highlighted leading column is max, which is not generally available.

The table's provenance also varies: Meta says it selects the highest comparable result available from its own evaluation, an official leaderboard, or a provider's self-reported number. It calls runs of competing proprietary models “best-effort.” The numbers are useful, but the max column is preview evidence—not a clean purchasing test of products available under the same conditions.

The version developers can use is still strong. Artificial Analysis gives Muse Spark 1.3 xhigh a score of 61 on its Intelligence Index. The page lists a public API price of $1.25 per million input tokens and $4.25 per million output tokens, and measures about 191 output tokens per second. Its separate max page reports 62, just one point higher, while labeling that mode “not publicly available” and showing no public speed or price.

A one-point composite gap does not mean the two settings behave alike on every task. It does mean the release is substantive: xhigh is a competitive, measurable product even without the unreleased mode leading Meta's chart.

Efficiency still depends on the workload. Meta says its engineers saw about 20% fewer tool calls and 25% fewer tokens than Spark 1.2 on coding tasks. Artificial Analysis saw xhigh emit 100 million output tokens across its broader index, against a 71 million median for the comparison class; max emitted 120 million. These results do not contradict each other because they measure different tasks and reasoning settings. They do show why “25% fewer tokens” should not be carried from one test suite into a general cost forecast.

Open weights are still a promise. Meta lists a Muse Spark open-weights release on its roadmap, but the launch includes no checkpoint, license, parameter count, or date. “Coming soon” is not an artifact developers can inspect or run.

The useful watch items are concrete: when max reasoning becomes generally available, what its safety gate actually tests, whether same-setting evaluations reproduce the gains, and what Meta eventually publishes under the open-weights label. Until then, compare the xhigh model you can use—not the max column you can only see.

Correction, Sep. 3. An earlier version said the headline comparison was not a same-setting test. Meta's full PDF table includes both 1.3 max and 1.3 xhigh, so it does allow an xhigh-to-xhigh comparison with 1.2. The unavailable max mode remains the highlighted leading column. I also corrected the Bluesky thread.

Source graph: Semble source collection