Salesforce announced a benchmark for evaluating large language models (LLMs) on customer relationship management (CRM) work on June 18, 2024. It compares models across sales and service tasks using four decision dimensions: accuracy, cost, speed, and trust and safety. Salesforce called it the world’s first LLM benchmark for CRM; that is the company’s description of its launch.
What Salesforce’s CRM benchmark is
The benchmark is a framework for comparing how well LLMs handle work inside CRM systems, rather than judging them only on general-purpose or academic tests. Its initial focus is common sales and customer-service tasks, including prospecting, lead nurturing, sales-opportunity summaries, and service-case summaries.
Salesforce said the results are available through an interactive Tableau dashboard and a Hugging Face leaderboard. Coverage and rankings may change: the company said it planned to add use-case scenarios and later include fine-tuned models.
What it measures
The framework brings four considerations into one view. A model that performs well on one dimension may not be the best operational choice if it is expensive, slow, or unsuitable under an organization’s safety requirements.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
| Dimension | What it covers |
|---|---|
| Accuracy | Factuality, completeness, conciseness, and following instructions. |
| Cost | Low, medium, and high categories based on percentiles; the supplied description does not specify the underlying thresholds. |
| Speed | Responsiveness and processing efficiency. |
| Trust and safety | Protection of sensitive customer data, privacy, security, bias, and toxicity. |
For model selection, accuracy is most useful when examined by task and by its component measures. A concise summary, for example, is not necessarily a complete or factually reliable one. Safety is also broader than avoiding toxic text: the stated evaluation includes sensitive-data protection, privacy, security, and bias.
How Salesforce built the initial evaluation
Salesforce AI Research identified 11 common CRM use cases across sales and service, created standard prompt templates, and grounded those prompts with real CRM examples. The initial study ran the tasks against 15 LLMs.
Rank #2
Salesforce employees and external customers or other practitioners assessed the outputs, with automated LLM judges used to scale evaluation. That combination offers practitioner input alongside automated review, but readers should not treat an automated judge’s score as equivalent to a human assessment. The benchmark description does not establish that every output or every metric was evaluated exclusively by human reviewers.
How to use the leaderboard when choosing a model
- Start with the workflow. Identify the specific sales or service task you need the model to perform, such as lead nurturing or summarizing a service case. A broad overall ranking may obscure differences between use cases.
- Compare the accuracy components. Check factuality, completeness, conciseness, and instruction-following for the relevant task, rather than relying on a single accuracy label.
- Check operating trade-offs. Compare cost categories and speed alongside quality. The published description characterizes cost in percentile-based low, medium, and high bands, so do not infer an exact price from a category alone.
- Review the safety fit. Consider the benchmark’s coverage of customer-data protection, privacy, security, bias, and toxicity against your organization’s own policies and risk review.
- Check how the result was judged and what is covered. Note the role of practitioner review and automated judging, and confirm that the leaderboard’s task coverage and model set still match the decision you are making.
- Validate in your own pilot. The benchmark can inform shortlisting, pilot design, and governance reviews; it does not by itself establish how a model will perform with your data, configuration, or controls.
What the benchmark can—and cannot—tell a buyer
Its value is that it frames model choice around CRM work and business constraints, combining task quality with cost, responsiveness, and safety considerations. That makes it more directly relevant to a sales or service pilot than a general benchmark that does not represent those workflows.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
It is a comparison aid, not a guarantee of production performance or a substitute for an organization’s own testing and governance. Results depend on the benchmark’s selected use cases, prompts, models, and judging approach; Salesforce also said coverage could evolve. The initial study’s 15-model scope should therefore be understood as the set evaluated in that study, not as a complete or permanent inventory of available models.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




