Salesforce this week introduced its first purpose-built CRM reasoning model, called Koa, at Dreamforce 2026 in San Francisco — a product that embodies the company’s strategic conviction that enterprise AI needs domain-specific models rather than general-purpose large language models. The problem is that Koa’s most specific performance claim, that it “matches or exceeds leading models on complex CRM tasks with three times fewer errors,” comes from a benchmark Salesforce designed and administered itself. No independent third party has verified that figure, and a separate Bloomberg investigation published before Dreamforce found enterprise customers reporting that Agentforce, the AI agent platform Koa is designed to power, is producing outcomes that materially miss what Salesforce’s marketing promised. For any organization evaluating whether to add Koa to its AI roadmap, those two facts belong together in the same sentence.
Alongside Koa, Salesforce launched AIforce at Dreamforce 2026, a new architectural layer designed to make Salesforce’s CRM data, permissions, and business logic accessible from inside third-party AI environments — including Anthropic’s Claude, Google Gemini Enterprise, and Amazon Q — without employees opening a Salesforce tab. The announcement represents something the company has not said plainly in its marketing but that its own president of applications, Patrick Stokes, articulated more directly: “The value of Salesforce is in the data and the metadata. What we are doing now is exposing it to a new UI.” That statement, made at the world’s largest CRM conference, is an unusually candid acknowledgment that Salesforce’s competitive future no longer rests on the interface it spent 25 years building.
What Koa Is — and How It Actually Works
Koa is a reasoning model built specifically for CRM tasks, developed by post-training NVIDIA’s Nemotron 3 Super base model using a proprietary synthetic dataset. The training set was derived from what Salesforce describes as approximately 27 years of CRM domain knowledge, generating synthetic scenarios across more than 14 industries. Crucially, Salesforce used no real customer data in the process — a deliberate choice made for both privacy and legal reasons.
The underlying Nemotron 3 Super architecture uses a hybrid Mixture of Experts design. MoE models do not activate all of their parameters for every inference request. Instead, a learned routing mechanism selects a small subset of specialized subnetworks — called “experts” — for each input. In practice, this means a model with very large total parameter capacity can run at a fraction of the compute cost of a dense model, because most of its parameters sit idle during any given request. For Salesforce, this matters economically: Agentforce’s base pricing of $2 per conversation only makes sense if the underlying inference is cheap.
Koa is also built with a one-million-token context window, which enables an agent to reason across an entire customer history, a long document chain, or a complex multi-step workflow without losing the thread. The model is specifically optimized for the CRM failure modes that most damage enterprise trust: updating opportunity stages incorrectly, misrouting customer support cases, and scheduling follow-ups for the wrong contacts.
NVIDIA CEO Jensen Huang, appearing on the Dreamforce keynote stage alongside Salesforce Chair and CEO Marc Benioff, described the partnership as an example of turning “decades of enterprise expertise into specialized AI.” The key phrase is “specialized” — and that word carries a specific technical bet. Salesforce and NVIDIA are arguing that a model post-trained on CRM-specific synthetic scenarios will outperform a general-purpose frontier model on CRM tasks, even though the frontier model has vastly more total parameters and has seen far more diverse training data. That bet may prove correct. It may also be the kind of claim that looks compelling in a keynote and proves harder to defend in production.
The Benchmark Question Salesforce Has Not Answered
Koa’s headline performance claim is that it “matches or exceeds leading model performance” on Salesforce’s internal CRM benchmark while producing three times fewer errors on specific CRM tasks. That internal benchmark is Salesforce’s own creation, designed by the same organization that stands to benefit from strong Koa performance. No independent auditor has published a verification of those results.
This matters because the AI industry has a documented problem with benchmark reliability. A UC Berkeley Research, Development, and Innovation team published findings in April 2026 that researchers had successfully manipulated eight industry-standard AI agent benchmarks to near-perfect scores without solving a single real task. Apple’s GSM-Symbolic research found performance drops of up to 65 percent when models that scored highly on one math benchmark were tested on minor variations of the same problems. The international AI safety community’s 2026 report concluded directly that “performance on pre-deployment tests does not reliably predict real-world utility or risk.”
None of this proves Koa’s benchmark is gamed or wrong. It proves that a single vendor-administered benchmark on a vendor-developed evaluation set is insufficient evidence to make a purchasing decision. Koa is currently in pilots with 1-800Accountant, Baxter Credit Union, Formula 1, UChicago Medicine, and Xero, with U.S. general availability not expected until winter 2026. By the time those pilots produce deployable evidence, organizations evaluating Koa today will be working from the company’s word.
AIforce and the Death of the Salesforce Tab
The more architecturally significant announcement at Dreamforce 2026 may not be Koa itself but AIforce, the new interface layer that Salesforce introduced to sit above its existing Data 360, Customer 360, and Agentforce layers. AIforce’s purpose is to make Salesforce’s entire stack accessible from any AI environment that supports the Model Context Protocol (MCP), the open standard Anthropic introduced in November 2024 to solve what had been an “N×M” integration problem in enterprise AI.
Before MCP, connecting an AI assistant to an enterprise’s data required building custom connectors for every AI-to-data- — comparable, in the industry’s own analogy, to what USB-C did for physical device connectivity. Salesforce’s Headless Toolkit now exposes more than 60 MCP tools, more than 30 coding skills, more than 4,000 APIs, and more than 220 command-line interface commands to external builders
The three initial AIforce products are Claudeforce, Slackforce, and Agentforce Coworker. Claudeforce places a prebuilt MCP server carrying 37 ready-made sales skills directly inside Anthropic’s Claude, covering tasks from prospecting to pipeline management. Anthropic CEO Dario Amodei appeared onstage at Dreamforce to discuss the partnership; he noted that even with AI frozen at current capabilities, organizations are using perhaps 5 to 10 percent of the value the technology could provide. Slackforce brings Salesforce context directly into Slack conversations, allowing the Slackbot personal assistant to reason over both Slack conversation history and governed Salesforce information simultaneously. Agentforce Coworker operates within Lightning and Salesforce’s existing permissions framework; the company reported that 100,000 users activated it during its first 35 days — a figure Salesforce has not supported with an independent reference.
What AIforce collectively represents is a strategic pivot from UI competition to data-layer competition. Benioff captured the intended framing during the keynote: “AI models show what’s probable. Core systems determine what’s permitted.” The flip side of that statement is that Salesforce’s legacy competitive advantage — the best CRM dashboard — is increasingly beside the point if AI agents can access its data directly.
What the Bloomberg Finding Means for Enterprise Buyers
The timing of Dreamforce 2026’s ROI-themed Day 2 keynote — formally titled “From Fast Start to Real ROI” — is not coincidental. Bloomberg published an investigation before the conference documenting that enterprise customers are reporting Agentforce outcomes that materially miss the company’s marketed expectations. A Goldman Sachs analysis cited alongside those findings noted that agent token economics are shifting rapidly, creating unstable ROI calculations across different deployment contexts.
Independent reviews are consistent with the Bloomberg picture. A 2026 technical review of Agentforce found implementation timelines of four to six weeks for single-use-case deployments and eight to sixteen weeks for complex rollouts, with Year 1 total costs reaching $270,000 to $540,000 when Data Cloud licensing, implementation services, and training are factored in. Only 27 percent of commerce organizations in Salesforce’s own State of Commerce survey reported having fully unified customer data across sales, service, marketing, and commerce. Agentforce agents require clean, unified data to function correctly; an agent operating on duplicate or conflicting customer records produces confident, incorrect actions — not useful automation.
Futurum Research data published in 2026 placed Salesforce’s 34.1 percent CRM market share in an $85.4 billion market. Its position as an AI platform vendor is early-stage and contested, with Microsoft Copilot Studio (160,000 organizations, 400,000-plus custom agents) and ServiceNow both building competing agent ecosystems.
Day Two and the Siemens Case Study
The most detailed customer evidence at Dreamforce 2026 came from Siemens, whose President and CEO Roland Busch joined Benioff at YBCA on Day 2 to discuss a two-agent deployment inside Sales Cloud. Siemens reports it had been receiving more than 2,500 unqualified leads per month across its approximately 18,000 sellers in 132 countries. Two AI agents now qualify prospects and route stronger opportunities to human sellers, with Siemens reporting 100 percent lead engagement. Busch described the integration as compressing some industrial processes from weeks to hours.
Siemens is a meaningful proof point — a large, complex enterprise with a specific, measurable use case and a senior executive willing to go on a public stage. But it is a single customer in a single sales workflow, and Salesforce has not disclosed similar outcome data at scale. The Agentforce keynote’s own framing — “from fast start to real ROI” — implicitly acknowledges that most organizations are still in the first part of that journey.
Separately, Salesforce and Google Cloud announced that Hyperforce on Google Cloud is already handling live production customer traffic, with North American general availability scheduled for November 2026. On the AWS side, Salesforce business context now flows into Amazon Q through the headless architecture, and Amazon Bedrock is becoming a broader model choice
Day Three: What Is Still Ahead Today
As of this morning, Dreamforce’s final day is underway but its marquee afternoon sessions have not yet occurred. The Slack Keynote is scheduled to begin at 10 a.m. PT (1 p.m. ET), positioning Slack as the entry point through which employees will interact with agents in daily work. The Admin Keynote, which will cover Flow, Prompt Builder, and the admin toolkit for agent-native organizations, is scheduled for 10:30 a.m. PT.
True to the Core — the beloved unscripted Q&A format in which Salesforce co-founder Parker Harris and other product leaders take live questions from the floor — is also scheduled for this afternoon. The format, which has been a Dreamforce fixture for years, historically surfaces candid answers about platform limitations and roadmap items that the polished keynotes do not. Reporters covering the event will be watching it closely for any commentary on the ROI gap Bloomberg documented.
A final YBCA conversation between Benioff and Travis Kalanick is scheduled for 12:30 p.m. PT (3:30 p.m. ET). The event closes tonight with a Comedy Hour featuring Jim Gaffigan.
The Salesforce+ virtual program continues through Friday, September 18.
What Koa’s Architecture Reveals About Salesforce’s Competitive Strategy
Koa’s existence is itself a statement about where Salesforce believes the enterprise AI market is heading. For the past two years, Agentforce’s reasoning capability came from third-party models — OpenAI, Anthropic, Google — routed through the Atlas Reasoning Engine. Koa represents Salesforce deciding that relying on third-party models creates a long-term dependency risk and that the CRM domain is specific enough to justify a proprietary model.
The argument has precedent. NVIDIA’s Nemotron work has shown that domain-specific fine-tuning can produce models that punch above their weight on targeted tasks. A model with a fraction of the parameter count of GPT-5 or Claude Sonnet 5 can still outperform them on a specific, well-defined domain if the training data quality is high and the benchmark design is rigorous.
The caveat is exactly there: if the benchmark design is rigorous. Salesforce has not published Koa’s benchmark methodology, has not invited independent researchers to replicate the results, and has not released Koa under an open license that would allow third-party evaluation. Until one of those things happens, the “three times fewer errors” claim sits in the same category as every other vendor-administered AI performance number: plausible, interesting, and unverified.
How Does Koa Differ From Microsoft Copilot’s Approach?
Microsoft Copilot Studio, Salesforce’s primary enterprise AI agent competitor, has taken a different architectural path. Rather than building a domain-specific model, Microsoft has integrated GPT-5 and optionally Claude into Copilot Studio’s agent framework, leveraging the raw capability of general-purpose frontier models while embedding them in its existing enterprise permissions infrastructure. With 160,000 organizations running more than 400,000 custom agents, Microsoft’s installed base advantage is substantial.
Salesforce’s counter-argument — implicit in Koa’s existence — is that general frontier models, however capable, do not understand CRM data structures, CRM-specific failure modes, or enterprise governance constraints without extensive prompt engineering. Koa is designed to need less of that scaffolding because the CRM knowledge is baked in. Whether that tradeoff produces better outcomes for enterprise buyers than a well-prompted frontier model is a question that Koa’s winter 2026 general availability will begin to answer.
What Salesforce Did Not Say at Dreamforce 2026
Three significant gaps characterized what Dreamforce 2026’s keynotes left unaddressed. First, Salesforce disclosed no pricing for AIforce. The company introduced an architectural layer with 60-plus MCP tools and 4,000-plus APIs and declined to say what it costs. For enterprise architecture teams evaluating whether to build on the Headless Toolkit, that is a material omission.
Second, Salesforce disclosed no general availability date for Koa outside of named pilots in the United States, with “winter 2026” as the only guidance. Organizations in regulated industries — healthcare, financial services — may have their own compliance timelines, and Salesforce provided no equivalent EU or APAC availability guidance.
Third, and most consequentially, Salesforce provided no detailed framework for multi-agent governance — how organizations should manage liability, auditability, and decision ownership when multiple AI agents from different vendors are making decisions in concert. The keynotes described the capability. The governance architecture, which will determine whether enterprises can deploy that capability safely, was not addressed.
What the Education Grants Signal About Salesforce’s San Francisco Relationship
Dreamforce 2026 included a pledge of $27 million in education grants, with $18 million directed to public school districts in San Francisco, Oakland, New York City, Chicago, and Indianapolis. The gift brings Salesforce’s reported cumulative contribution to public schools above $180 million, with more than $156 million directed to Bay Area schools. Oakland schools that received earlier Salesforce investments saw an eight-point improvement in standardized math scores; San Francisco middle schools recorded nearly 8 percent growth in math proficiency since 2022.
Separately, Salesforce launched the Missionforce Military Fellowship, a 12-week SkillBridge program placing transitioning service members alongside Missionforce Customer Success teams to build AI and CRM skills for civilian careers.
Dreamfest — the benefit concert supporting UCSF Benioff Children’s Hospitals — featured Usher and Gwen Stefani, who gave a preview performance at the keynote during Tuesday’s opening.
Frequently Asked Questions
What is Salesforce Koa and why does it matter?
Koa is Salesforce’s first purpose-built CRM reasoning model, developed by post-training NVIDIA’s Nemotron 3 Super on a synthetic dataset of CRM scenarios across more than 14 industries. It uses a Mixture of Experts architecture, which enables large model capacity at lower inference cost — important for Agentforce’s $2-per-conversation pricing to be economically viable at scale. Its significance depends on whether you believe enterprise AI needs domain-specific models (Salesforce’s argument) or whether well-prompted frontier models from OpenAI, Google, or Anthropic can do the same job with less vendor lock-in (the competing argument). What is not yet resolved is whether Koa’s benchmark claims — produced by Salesforce on Salesforce’s own benchmark — will hold up to independent scrutiny.
Why can’t organizations just use a general LLM like Claude or GPT-5 for CRM tasks instead of Koa?
They can, and many do — Microsoft Copilot Studio supports GPT-5 and optionally Claude in its enterprise agent framework. Salesforce’s argument for Koa is that general frontier models require extensive prompt engineering to understand CRM data structures, CRM-specific failure modes (like incorrect opportunity stage updates), and enterprise governance constraints. Koa is designed to have those constraints baked into its training rather than scaffolded through prompts. Whether that produces meaningfully better outcomes than a well-configured frontier model is the question that Koa’s winter 2026 general availability pilots are designed to answer.
What is AIforce, and what does it mean that Salesforce is calling its own UI obsolete?
AIforce is a new layer in Salesforce’s architecture that exposes the company’s CRM data, business logic, and permissions to any AI environment that supports the Model Context Protocol — including Claude, Gemini, and Amazon Q. Patrick Stokes, Salesforce’s president of applications, put it directly: the value is in the data and metadata, not the interface. What that means in practice is that Salesforce is competing less as a dashboard company and more as a data-infrastructure company. Organizations that have spent years building Salesforce data quality, permissions, and metadata potentially have a more valuable asset than they realized — but that value is now accessible to any MCP-compatible AI, which changes the competitive dynamics for both Salesforce and its enterprise customers.
Does Agentforce actually deliver ROI, given the Bloomberg investigation?
The evidence is mixed. Bloomberg found enterprise customers reporting outcomes that materially missed Salesforce’s marketing claims. Independent technical reviews find implementation costs of $270,000 to $540,000 in Year 1, with timelines of four to sixteen weeks per use case depending on complexity. The Siemens deployment, presented at Dreamforce, shows genuine production outcomes in lead qualification. But Siemens is a single case, and Salesforce has not published broader adoption metrics with outcome data at scale. The most honest framing: Agentforce can deliver ROI in specific, well-scoped use cases with clean underlying data — and struggles where data quality is low or the use case spans too many workflows at once.
ⓒ 2026 TECHTIMES.com All rights reserved. Do not reproduce without permission.
