Salesforce has introduced Koa, its first CRM reasoning model for Agentforce, as a specialized system for multi-step business work rather than a general chatbot. Koa is currently available to select pilot customers, while Salesforce expects general availability in U.S. regions in Winter 2026. The model is built by post-training NVIDIA’s open-weight Nemotron-3-Super-120B with Group Relative Policy Optimization, or GRPO.
Salesforce’s CRM reasoning model for Agentforce
Koa is designed to work through CRM tasks that require several decisions and tool calls. The examples supplied by Salesforce include updating opportunities, routing cases, scheduling follow-ups, selecting the right business tool, and completing multi-turn workflows.
That distinction matters. A conventional chatbot can produce a plausible reply; Koa is designed to resolve a workflow by choosing actions, changing the relevant system state, and reaching the requested outcome. Salesforce describes the model as an enterprise language model for agentic tool use, with Agentforce providing the environment where those actions take place.
How Koa fits into Agentforce
Salesforce says Koa can be selected in Data Cloud’s generative-model catalog, at the Agentforce organization level, and for individual agents and sub-agents. Salesforce also says it runs Koa’s model weights, post-training and inference inside its own infrastructure and trust boundary.
The current rollout is limited to select pilot customers. Salesforce expects U.S. general availability in Winter 2026, so the product is not yet a broadly available CRM model.
The model is also designed to operate with structured Salesforce workflows. Agent Script specifications define routers, sub-agents, typed actions, tool scopes and workflow instructions, giving the training system a more concrete target than text quality alone.
Nemotron-3-Super-120B and GRPO: how Koa was trained
Koa’s foundation is Nemotron-3-Super-120B, NVIDIA’s open-weight model with approximately 120 billion parameters. Salesforce then post-trained that base model with GRPO, a reinforcement-learning method that compares candidate action sequences and reinforces the trajectories that complete simulated tasks correctly.
The training pipeline turns Agent Script workflow definitions into simulated, multi-turn CRM scenarios. Simulated personas interact with the system, tools execute within the environment, and the reward system considers grounded answers, successful tool calls, correct state changes and workflow completion. In other words, Koa is trained to do the CRM job—not merely to sound as if it understands the job.
Salesforce says the scenarios cover more than 14 industries, including financial services, manufacturing, healthcare and travel. It also says the training used public and synthetically generated interactions and did not use customer data.
What the published performance tests actually show
Koa scored 0.86 on the overall CRM Bench evaluation described in Salesforce’s technical report. In that same comparison, Opus-4.8 scored 0.87, GPT-4.1 scored 0.81 and GPT-5.5 scored 0.90.
| Model | CRM Bench overall score |
| Salesforce Koa | 0.86 |
| Opus-4.8 | 0.87 |
| GPT-4.1 | 0.81 |
| GPT-5.5 | 0.90 |
The numbers place Koa close to Opus-4.8, above GPT-4.1 and below GPT-5.5 on this overall score. They do not support a claim that Koa is the strongest model overall. The evaluation was produced by the Koa authors, and Salesforce separately claims that Koa generates three times fewer errors than leading models on CRM actions. That claim uses a different metric from the overall CRM Bench score.
The technical report also gives Koa a function-call accuracy of 0.77, compared with 0.71 for its Nemotron base. On BFCL, Koa scored 66.63%, versus 64.73% for the base model and 53.96% for GPT-4.1. On Tau2Bench, Koa scored 69.41, compared with 68.64 for its base and 54.51 for GPT-4.1.
The practical takeaway is narrower—and more useful—than a general leaderboard victory: Salesforce is optimizing a model for CRM actions, tool selection and workflow completion, then measuring it in those conditions.
Availability, privacy and controlled deployment
Salesforce announced Koa with NVIDIA at Dreamforce on September 15, 2026. The company says the model is already used internally, including in an employee assistant in Slack, while select customers can access the pilot through Agentforce.
For organizations evaluating an agent with permission to update records or trigger business tools, the deployment boundary is part of the product story. Salesforce says it controls Koa’s weights and runs inference within Salesforce infrastructure. Its stated training policy is also specific: public and synthetic interactions were used, and customer data was not used to train Koa.
Salesforce and NVIDIA have additionally connected the model work with Missionforce, extending NVIDIA-based deployment into government and regulated environments, including private clouds and air-gapped networks.
Specialization before general superiority
Koa’s pitch is not that a CRM-trained model will replace every general-purpose frontier model. The technical results do not show that. Its more focused bet is that a model trained around business objects, tool permissions, typed actions and workflow state can be more useful for enterprise operations than a broadly capable model asked to improvise its way through a CRM.
That is why the important test will be operational: whether Koa can select the right tool, preserve context across several turns and make the correct state change inside a real organization’s Agentforce setup. Salesforce’s pilot rollout is the next stage of that test, with broader U.S. availability planned for Winter 2026.