The reasoning stack is the SKU

Yesterday’s Agentforce conversation was about named job packs: which role gets a multi-week goal, which agent sits in the org chart. Today’s Dreamforce drop changes the layer underneath.

On September 15, 2026, Salesforce and NVIDIA announced Koa, Salesforce’s first CRM reasoning model for Agentforce. It is post-trained from NVIDIA Nemotron 3 Super on synthetic scenarios modeled on nearly three decades of CRM deployment knowledge. Salesforce says it controls the weights and runs inference inside its own trust boundary. Select customers can pilot it now. U.S. general availability is expected in winter 2026. There is no public list price.

That is not another job title in the Agentforce catalog. That is a reasoning SKU. If you still evaluate Agentforce only as “which named agent do we buy,” you will miss the decision that actually moves spend and risk this quarter: which model stack sits under that role, where the tokens run, and whether a pilot seat is enough to rewrite your frontier-model plan before winter GA.

Evidence bar up front. I do not have hands-on pilot access. Claims below are grounded in Salesforce’s September 15 press release, the Salesforce Agentforce & AI Research arXiv paper (2609.15066), and independent coverage from SiliconANGLE, IT Pro, TechCrunch, and Pulse2. A Salesforce newsroom story URL that indexed as `/news/stories/koa-reasoning-model/` returned HTTP 404 during research, so it is not cited. Treat the vendor “three times fewer errors” line as Salesforce’s CRM benchmark, not an independent audit. The research paper’s published tables are more careful than the keynote slide.

What it is

Koa is a specialized language model for agentic CRM work inside Agentforce. Salesforce’s press release frames it as purpose-built to reason through complex, multistep workflows and to pick the right tools along the way: lead generation, opportunity qualification, case routing, follow-ups, and similar Customer 360 actions.

The base is NVIDIA Nemotron 3 Super (the arXiv paper specifies the open-weight Nemotron-3-Super-120B). Post-training uses synthetic data only. Salesforce is explicit: no customer CRM records in the training corpus. Scenarios simulate enterprise workflows across more than 14 industries (manufacturing, financial services, healthcare, travel, and others), pairing personas with task sequences and tool calls. Training stack language in the PR: supervised fine-tuning plus reinforcement learning with Group Relative Policy Optimization (GRPO), using NVIDIA NeMo RL, NeMo Gym, and NeMo AutoModel.

The research paper adds the operator-relevant mechanism. Enterprise tasks are built from Agent Script specifications (Salesforce’s declarative language for Agentforce agents). A simulation pipeline expands those specs into persona-conditioned multi-turn rollouts with rewards grounded in successful tool use. In plain English: the same agent-authoring artifacts that configure an Agentforce agent also structure how the model gets post-trained. That is a tighter loop than “fine-tune on a pile of tickets and hope.”

Inference posture is the commercial hook. Salesforce controls the weights and runs Koa on its infrastructure so customer data does not leave the trust boundary during inference. TechCrunch’s Julie Bort quotes Jayesh Govindarajan (EVP of Salesforce AI): before Koa, long multistep reasoning often routed through Agentforce’s AI gateway to frontier providers. Koa is meant to be a Salesforce-hosted alternative for that class of work, not a replacement for every Anthropic or OpenAI path Salesforce still sells.

Related track, not the same SKU: Nemotron-based models are also headed into Missionforce for regulated and air-gapped deployments, with post-trained models for Missionforce Operations slated for select customers in October 2026 per the PR.

What changed

GTM operators reviewing CRM workflow notes at a clean modern desk

Three shifts matter for GTM and RevOps operators. Only one of them is “Salesforce trained a model.”

First, the Agentforce buy is now two-axis. Named job packs (sales development, service, etc.) still describe *what role* you are automating. Koa describes *which reasoning stack* can sit under that role when the workflow is long, tool-heavy, and CRM-shaped. SiliconANGLE’s Mike Wheatley puts Kumar’s framing cleanly: domain-specific post-training that adds job-specific intelligence on top of a strong base, without training on customer data.

Second, sovereignty and provenance moved from slideware to a concrete option. TechCrunch reports Govindarajan’s rationale for waiting on Nemotron: a sovereign American pre-trained base with clearer data provenance than some open-weight alternatives. Whether you buy that geopolitics or not, the operator point is simpler. You can now ask Procurement and Security a yes/no question they could not ask last week: do we want CRM-agent reasoning to stay on a Salesforce-owned, Salesforce-hosted model for selected workloads?

Third, the benchmark story forked. The press release and SiliconANGLE both relay Salesforce’s CRM benchmark claim: match or exceed leading models on CRM actions with three times fewer errors on tasks like updating opportunities, routing cases, or scheduling follow-ups. The arXiv paper publishes a different, more granular picture. On Tau2Bench, Koa edges its Nemotron base (69.41 vs 68.64 weighted) and beats GPT-4.1 by a wide margin. On BFCL it improves over the base. On CRM Bench it scores 0.86 overall, near Claude Opus 4.8 (0.87) and above GPT-4.1 (0.81) and the base (0.84), with clearer function-call gains. The paper is blunt that Koa remains below the strongest frontier models. If your vendor deck only shows “3x fewer errors,” ask which table they mean.

What works (on paper)

Honest specialization thesis. Building a CRM reasoning model from synthetic workflow specs, not scraped customer orgs, is the right product story for a multi-tenant CRM vendor. IT Pro and the PR both stress synthetic scenarios mapped to tool-call sequences. That will not stop every privacy review, but it is a cleaner answer than “we trained on your industry somehow.”

Trust-boundary inference as a real differentiator. For regulated buyers and for CISOs who hate gateway hops to third-party frontier APIs, a Salesforce-hosted reasoning option is a concrete control. Pulse2 and the PR both hammer weights-plus-inference inside Salesforce infrastructure.

Pilot roster that looks like production pain, not logo theater. Named pilots in the PR and IT Pro: 1-800Accountant, Baxter Credit Union, Engine, Formula 1, UChicago Medicine, and Xero. Ryan Teeples (1-800Accountant) talks tax rules, documents, and step-by-step tool use. That is the shape of work where generic chat models burn tokens and still miss the next CRM action.

Internal dogfood signal. Salesforce says Koa already powers an internal Slack agent for employee tasks. Dogfood is not a customer SLA. It is still better than a model that only exists in a keynote.

Research transparency above the average enterprise AI launch. Shipping an arXiv report with tables vs Opus 4.8, GPT-5.5, GPT-4.1, and the Nemotron base is unusual for a Dreamforce SKU. Use it.

What breaks (or stays opaque)

Enterprise ops team in a bright meeting room with laptops and a blank whiteboard

Pilot-only is not a procurement plan. Availability line from Salesforce: select Agentforce pilot customers now; GA expected winter 2026 in U.S. regions. If your FY planning assumes GA seats in October, you are inventing a date Salesforce did not publish.

No public price. None of the primary or secondary sources reviewed for this piece list a Koa price, credit multiplier, or Agentforce add-on SKU. You cannot model TCO from the homepage. You model it after your AE confirms whether Koa is bundled, metered, or gated behind a higher Agentforce tier.

Vendor benchmark vs paper tables. “Three times fewer errors” is Salesforce’s CRM benchmark language. It is not replicated as a single headline number in the arXiv results section. Operators who paste the 3x claim into a board deck without citing “Salesforce benchmark, unaudited by us” are doing marketing, not diligence.

Still below the strongest frontier on the paper’s own charts. If your hardest Agentforce workflows already depend on Opus-class or GPT-5.x reasoning through the gateway, Koa is a candidate routing target, not an automatic rip-and-replace. TechCrunch is clear Salesforce is not abandoning Anthropic or OpenAI paths.

Synthetic training is not your org’s process graph. Fourteen industries of simulated workflows are not your opportunity stages, your entitlement rules, or your messy custom objects. Expect configuration, evaluation, and failure analysis on *your* Agent Script and tool scopes. The paper’s strength is the spec-to-reward loop. Your strength has to be the eval set.

I cannot tell you latency, tool-miss rate, or token cost on your tenants. No hands-on means no production diary. Press quotes and vendor benches are not a pilot readout.

Who it is for

Strong fit (as described): Agentforce customers already deep in Salesforce data and permissions who need multistep CRM agents (case routing, opportunity hygiene, lead qualification, follow-up orchestration) and who prefer reasoning tokens to stay inside Salesforce’s trust boundary. Pilots in accounting, credit union service, travel ops, and healthcare ops map to that pattern.

Conditional fit: Teams that currently route hard Agentforce reasoning to Claude or ChatGPT through the gateway and want a Salesforce-native alternative for a slice of traffic. Worth a controlled A/B on a single workflow family. Not a weekend self-serve toggle.

Weak fit for now: Orgs not on Agentforce, teams that need GA and a published price before any architecture change, or shops whose real bottleneck is dirty CRM data and undefined processes. A sharper reasoning model will not invent clean stages. Missionforce buyers should track the October Nemotron line separately.

Pricing and limits

  • Access: select pilot customers in Agentforce now (Salesforce PR, corroborated by SiliconANGLE and Pulse2).
  • GA: expected winter 2026, U.S. regions (Salesforce PR).
  • List price: none published in the PR, arXiv paper, or the independent pieces reviewed here.
  • Training data claim: synthetic scenarios only; no customer data (PR, SiliconANGLE, TechCrunch, arXiv).
  • Inference: Salesforce-controlled weights and infrastructure; customer data should not cross the trust boundary during Koa inference (PR).
  • Missionforce Nemotron models: select customers October 2026 (PR).
  • Benchmark claims: Salesforce CRM benchmark “3x fewer errors” (vendor); arXiv tables show incremental gains over Nemotron base and wins vs GPT-4.1, still below top frontier models.

What to do this week

  1. Rewrite the Agentforce one-pager with two columns. Left: named roles / job packs. Right: reasoning stack options (frontier via gateway vs Koa vs other hosted models). If the right column is blank, your buy sheet is last week’s sheet.
  1. Ask Security three questions in writing. Does Koa inference stay entirely in Salesforce’s trust boundary for our instance and regions? What logging exists for tool calls and model routing? How does Zero Data Retention interact when some requests still fan out to Anthropic or OpenAI?
  1. Pick one multistep CRM workflow for a pilot eval, not a logo tour. Example shapes from the PR and pilots: opportunity update chains, case routing with policy checks, accountant-style document-plus-rules flows. Build a 20 to 50 case gold set with expected tool sequences. Score Koa vs your current gateway model on task completion, wrong-tool rate, and human edit distance. Ignore the 3x slide until that table exists.
  1. Do not change FY model spend on pilot vapor. If you are not in the select pilot, treat winter 2026 U.S. GA as the planning fence. If you are in pilot, time-box a 30-day readout with kill criteria before you renegotiate frontier commitments.
  1. Put the arXiv tables next to the 3x slide for Procurement and the AI council. Alignment between those artifacts is your honesty check.

Close

Named Agentforce agents tell you who is on the digital org chart. Koa tells you what kind of brain can sit in that seat without shipping every hard token to a frontier lab.

The operator mistake this week is collapsing those into one SKU conversation. The reasoning stack is now a separate line item in the Agentforce design, even when the invoice has not published a price. Pilot-only access is enough to start an eval. It is not enough to pretend your model/vendor plan already changed.

If you cannot say which workflows will stay on Claude or GPT, which will try Koa, and what evidence retires either choice, you do not have a Dreamforce takeaway. You have a press release.