Stop Scoring the Mode You Can’t Call

The slide deck still has two columns.

OpenAI. Anthropic. Maybe a polite footnote for Google Flash. Then someone on the buying committee pastes a screenshot of Meta's new chart — Muse Spark 1.3 sitting next to Claude Opus 5 and GPT-5.6 Sol on coding and agent work — and the room goes quiet for the wrong reason.

Because the chart is real enough to wreck a shortlist. And the mode that produces the headline numbers is not the one your API key can call this morning.

Meta just put a coding-and-agent model on the frontier shortlist at last month's API price — and the chart that sells it scores a max mode you still cannot buy.

What actually shipped this week

On September 2, Meta released Muse Spark 1.3 into Muse Code and the Meta Model API the same day, according to Meta's own model page at developer.meta.com and contemporaneous reporting. Bloomberg's September 2 interview with chief AI officer Alexandr Wang — carried September 3 by The Star — called it Meta's "biggest jump so far on model performance," aimed at coding and agentic work. Wang told Bloomberg the update is "competitive" with Anthropic's Claude Fable 5.1 and "better than" OpenAI's GPT-5.6 Sol on coding, while noting OpenAI's Astra was still about to land.

VentureBeat's September 3 read is the operator version of the same story. The shipping configuration uses Meta's previously available reasoning settings, including xhigh. Max reasoning — the column that carries Meta's strongest published scores — is still finishing additional safety testing and arrives "shortly." Artificial Analysis, cited by VentureBeat, scored the max preview at 62 on its Intelligence Index and the broadly available xhigh build at 61. That is not vapor. It is also not the chart on the launch graphic.

Same day, Google shipped Gemini 3.8 Flash at the same introductory price as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output through December 31, 2026, per Google's September 2 blog. Third Flash release in six weeks. Pitch: long-horizon coding and autonomous agents. Also a Cyber variant gated through Fairwind — a different essay.

Then September 3, OpenAI started rolling GPT-6 Astra to Trusted Access enterprises, with Plus, Pro, Business, Enterprise, and API access "coming in the coming days," per OpenAI's model docs. Price on the docs page: $10 input / $50 output per million tokens. Path to Astra on September 1 already said the model met OpenAI's Critical cybersecurity capability threshold and that advanced cyber access would stay limited. NBC News the same day confirmed the limited Thursday rollout and the safety designation.

Seventy-two hours. Three vendors. One shared move: the frontier is no longer a two-name RFP. Meta forced the shortlist open. Google kept pressure on the workhorse lane. OpenAI put the top of the stack behind preparedness and Trusted Access. If your procurement matrix still has two columns, it is already wrong.

The mechanism is the mode, not the model name

Here is the part most decks will skip.

Hands on a laptop showing a three-column comparison table beside a notebook labeled Shortlist.

Meta publishes an 11-row comparison that CellCog transcribed from Meta's September 2 materials: Muse Spark 1.3 (max) against Muse Spark 1.2, GPT-5.6 Sol, and Claude Opus 5. On Meta's own table, 1.3 leads or ties on DeepSWE v1.1 (75.4%), SWE-Atlas (59.4%), Terminal-Bench 2.1 (88.8%), and both long-context MRCR bands (98.5% and 98.1%). It sits within a point of Opus 5 on several agent and professional rows. Generation-over-generation jumps from 1.2 are large where 1.2 was weak: long-context retrieval, long-horizon coding, computer use.

Every one of those comparison-table numbers is max effort. Meta's availability language, as quoted across VentureBeat and CellCog, is careful: previously available reasoning modes ship today; max comes after more safety testing. VentureBeat even prints the gap on Meta's own disclosed pairs — GDPval-AA v2 at 1,754 Elo for max versus 1,709 for xhigh; OSWorld 2.0 at 66.9 versus 57.2; JobBench at 64.9 versus 61.2. Some rows are close or reverse. The point is not that Meta hid the shipping model. The point is that launch marketing leads with the configuration you cannot yet provision.

Pricing did not move. Wang told Bloomberg developers pay the same for 1.3 as for 1.2. VentureBeat and CellCog both list Standard at $1.25 input / $0.15 cached / $4.25 output per million tokens. There is a second door: a contributor tier at $0.10 input / $0.20 output, labeled for traffic Meta can use to improve its products. That is roughly a 95% haircut on output price in exchange for your prompts and completions becoming training material. Zuckerberg's "almost too cheap to meter" line, reported by VentureBeat, is not a price cut. It is a claim about what you can finish with the tokens you already buy — plus a discount tier that is a data-rights decision wearing a pricing costume.

Meta also claims behavioral efficiency: about 20% fewer tool calls and 25% fewer tokens than 1.2 on comparable coding tasks in its engineers' comparisons, plus asking clarifying questions and confirming before consequential actions. That is the right product instinct for agent loops. Treat it as a hypothesis to measure on your harness, not a finance-model input on day one. Artificial Analysis, again via VentureBeat, estimated xhigh at $0.55 per Intelligence Index task versus $0.40 for 1.2 — heavier agentic input consumption on their suite even at unchanged unit prices. "Cheap" gets slippery the moment the model runs as an agent.

What operators keep scoring wrong

Most eval packs still grade the press graphic.

Wrong exam.

If you score Meta's max column and then wire `muse-spark-1.3` at the effort level your account actually exposes, you built a fantasy bakeoff. Re-run at the mode you can call. Put the max column in a parking lot labeled "pending GA." Anything else is how you lose credibility with engineering two weeks after the board meeting.

Operators also confuse unit price with job price. Muse Spark Standard sits far under Opus 5 and Sol list rates on Meta's comparison set. That matters. It does not decide your bill. Reasoning effort, tool retries, context stuffing, and mid-task re-plans do. Google's own 3.8 Flash post admits the model "works harder" and may spend more tokens at higher effort. Meta claims the opposite direction versus 1.2 on coding. Both can be true on different workloads. Your traffic is the only adjudicator.

The contributor tier is the other silent failure mode. It will show up in a Slack thread at 1 a.m. as "why are we paying 20x for the same model id family?" Because one id sells your code and customer text back into Meta's training loop and the other does not. CellCog's reading of Meta's labels is blunt and correct: for anything you would not publish, Standard is the only door. The discount is the price of that rule being easy to forget under burn-rate panic.

Then there is the two-vendor hangover. Wang's "gemini who?" trash talk, reported by VentureBeat after Artificial Analysis posted results, is fun executive theater. The independent snapshot VentureBeat cites is tighter than the meme: Muse xhigh at 61 and $0.55 per task versus Gemini 3.8 Flash high at 59 and $0.58, with Google faster on throughput and cheaper on raw tokens during the promo window. That is not a knockout. That is a market where three workhorse options sit inside the noise of your eval harness. An RFP that pretends Meta is still "the open-weights company we ignore for production" is nostalgia, not analysis.

Second-order effects the launch posts will not say out loud

First, vendor charts became adversarial. CellCog caught the tell: Google and Meta both scored Claude Opus 5 on September 2 and disagreed on OSWorld and Terminal-Bench because harnesses, versions, and effort levels differ. DeepSWE matched because both pulled a public leaderboard. Same third-party model, different vendor theaters, same week. Trust generation-over-generation deltas from one lab. Treat cross-vendor rows as a range. Wait for independent indexes before you tattoo a ranking into a SOW.

Frosted glass conference room with a badge lock and a screen of abstract charts beyond the door.

Second, open-weights expectations are now a planning risk, not a brand promise. Bloomberg's Wang interview says Meta has not decided whether to release 1.3 weights and still plans weights for 1.2. VentureBeat notes the roadmap language drifted from a dated 1.2 promise toward a vaguer "Muse Spark open weights release." Teams that standardized on Llama for self-host economics cannot pencil "Spark weights next month" into capacity plans without a date, a size, and a license. Proprietary monthly cadence is the real Meta posture right now. Plan against that, not against the keynote memory of 2024.

Third, the preparedness stack is becoming a seat-license feature. OpenAI's Astra docs and Path to Astra make misalignment monitoring and cyber-gated access part of the product surface. Google's Fairwind Cyber SKU is the same shape in a different logo. Meta holding max behind safety testing is the quieter cousin of that pattern. Your "which model" decision now includes which capabilities are withheld, which monitors can pause a long agent run, and which approvals sit between your team and the mode in the chart. That is procurement, security, and platform engineering sharing one ticket.

Fourth, agent platforms that hard-bind to a single frontier name are about to look brittle. The model layer is rotating weekly — Fable 5.1 on September 1, Gemini 3.8 Flash and Muse Spark 1.3 on September 2, Astra on September 3. The durable question is whether roles, memory, permissions, and in-flight work survive the swap. If your stack dies when the preferred model id changes, you do not have an agent platform. You have a wrapper with a favorite.

What to do this week

Rebuild the shortlist with three production columns and one holding pen: OpenAI, Anthropic, Meta, plus Google Flash as the throughput/cost control. Drop any scoring sheet that cannot name the exact model id, effort level, and data-use tier under test.

Run Muse Spark 1.3 only at the reasoning mode your account can call today. Log tool-call counts and total tokens against the same coding and agent tasks you already use for Sol, Opus/Fable, and Gemini 3.8 Flash. If Meta's 20%/25% efficiency claim shows up on your harness, that is a real budget event. If it does not, you still learned something the press graphic cannot tell you.

Ban contributor-tier model ids from any environment that touches proprietary code, customer data, or regulated content. Put that rule in the same place you put "no production secrets in prompts." Make the expensive id the default and the cheap id an explicit exception with an owner.

For Astra, treat Trusted Access and Daybreak-class cyber pathways as separate products from the ChatGPT seat story. Read Path to Astra before you promise the board "we will just turn on Astra next week" for unrestricted agent loops. The docs already warn that monitors can pause or stop long-running work, including non-cyber tasks that look risky to the classifier.

If you buy agent platforms or "AI employees," ask one question in every demo: when Muse / Fable / Astra / Flash swaps underneath, what survives? Memory, roles, audit trail, permissions, and tasks in flight should be the answer. Model loyalty should not.

Close

The frontier did not get simpler this week. It got honest.

Meta earned a seat on the coding shortlist at an unchanged Standard price, with a shipping model that independent scores put in the frontier cluster and a launch chart that still leads with a mode behind the glass. Google kept the workhorse lane loud. OpenAI put the top rung behind preparedness and staged access.

Stop scoring the mode you cannot call. Buy the configuration you can provision, measure, and govern — or admit you are still shopping press releases.