CSAT, NPS, and CES benchmarks in 2026
The sources do not support one universal target. CSAT, NPS, and CES are not interchangeable scores. They answer different questions about different parts of the customer experience. Treating a score from one measure as a substitute for another can make a support operation look healthy while hiding friction, weak relationship sentiment, or a sampling problem.
This report uses current source material available on August 2, 2026. It reports published definitions and survey-method guidance, not a fabricated cross-industry benchmark. A benchmark is only comparable when its instrument, scale, population, timing, channel, response handling, and calculation rule are sufficiently similar to yours.
What CSAT, NPS, and CES each measure
Customer Satisfaction Score, or CSAT, is an in-the-moment assessment of satisfaction with a defined product, service, transaction, or interaction. Qualtrics gives a 1 to 5 example scale from very unsatisfied to very satisfied and describes a common top-box calculation using responses 4 and 5 divided by all survey responses. [4]
Net Promoter Score, or NPS, asks a relationship or advocacy question: how likely is the respondent to recommend the company, product, or service to a friend or colleague? Bain's published method uses a 0 to 10 scale, classifies 9 and 10 as promoters, 7 and 8 as passives, and 0 through 6 as detractors, then subtracts the percentage of detractors from the percentage of promoters. [1]
Customer Effort Score, or CES, asks how much effort the customer had to exert to complete a task, resolve an issue, or get an answer. Qualtrics describes it as a single-item measure that is commonly phrased on a very easy to very difficult scale. The direction and calculation must be stated because implementations can report an average, an easy-response share, or a net measure. [6]
| Measure | Core question | Unit commonly reported | Best-fit decision |
|---|---|---|---|
| CSAT | How satisfied was the customer with this specified experience? | Top-box percentage or average, with the rule disclosed | Diagnose a recent interaction or touchpoint |
| NPS | How likely is the customer to recommend the company, product, or service? | Promoter percentage minus detractor percentage | Track advocacy or relationship sentiment |
| CES | How easy or difficult was it to complete this task or resolve this issue? | Average, easy-response share, or net measure, with the rule disclosed | Find process friction at a touchpoint |
The measures can be used together, but a combined dashboard is not a combined score. A customer may be satisfied with a single interaction and still be unwilling to recommend the wider company. A customer may also receive a successful resolution after an effortful process. That is why the event being measured should be stored with every response.
The calculation is part of the benchmark
Two teams can use the same metric label and produce numbers that are not comparable. For CSAT, record the scale, the labels, the top-box rule, the denominator, and whether unanswered invitations are excluded. For NPS, preserve the 0 to 10 wording and the three category bands if you intend to compare with the Bain method. For CES, record the exact wording, direction, response scale, and whether the output is a mean, top-box, or net score.
| Metric | Reproducible calculation example | Disclosure needed |
|---|---|---|
| CSAT | Satisfied responses divided by all completed responses, multiplied by 100 | Satisfaction labels, scale, top-box rule, denominator |
| NPS | Percentage of 9 to 10 responses minus percentage of 0 to 6 responses | Recommendation wording, 0 to 10 scale, category bands |
| CES | The chosen average, easy-response percentage, or net formula | Wording direction, scale, and aggregation rule |
Do not convert a 1 to 7 CES into a 1 to 5 CES by simple arithmetic and call the result comparable. Do not compare a CSAT average with a CSAT top-box percentage. Do not compare a relationship NPS survey with a post-ticket NPS survey without labeling the different populations and moments.
Survey design can move the result
Survey wording is measurement. Pew Research Center notes that small wording changes can change responses, and recommends clear, specific questions with non-overlapping response options. The Center also advises keeping wording and context consistent when measuring change over time. [8]
Question order matters too. Earlier questions can provide a context that changes how respondents answer later questions. Pew describes randomizing question or response-option order as a way to distribute order effects across respondents, while noting that randomization does not make every effect disappear. [8]
Mode is another part of the instrument. A post-chat in-product prompt, an email survey, and an interviewer-assisted phone survey reach different people in different moments. Pew's questionnaire guidance treats wording, order, mode, translation, and trend comparability as connected design decisions. [9]
For a support benchmark, document at least:
- the invitation population and eligibility rules;
- the event that triggered the survey and the delay before sending it;
- the exact question wording and response labels;
- the scale direction and calculation rule;
- the channel, device, language, and collection mode;
- invitation, delivery, completion, and usable-response counts;
- exclusions, duplicate handling, weighting, and segmentation;
- the field dates and any operational incident that could affect responses.
This documentation is more useful than a benchmark label alone. It lets a reader decide whether the comparison is like-for-like.
Response bias and nonresponse are benchmark risks
A score describes respondents, not automatically every customer who was invited. Customers with unusually good or bad experiences may be more likely to answer. A survey shown immediately after a support interaction may overrepresent people still engaged with that interaction. A survey sent by a channel used by only part of the customer base can omit other experiences.
A response rate should still be reported, but it is not a complete quality verdict. AAPOR distinguishes response, cooperation, and completion rates and explains that even a high response rate does not by itself establish that nonresponse bias is absent. [10] AAPOR's standard definitions identify coverage, measurement, and nonresponse as separate parts of total survey error. [11]
Use that distinction in the report. Never publish a benchmark as if the percentage alone explains its reliability. Pair the score with its denominator, invitation base, response rate definition, sampling or census approach, and known limitations. If respondent mix differs from the invited population, show the mix or explain any weighting used.
How to use benchmarks responsibly in 2026
Use an external benchmark for orientation, not as an automatic target. First match the metric definition. Next match the population and event. Then compare the survey design and field period. Only after those checks should you decide whether the external number can frame an internal result.
An internal baseline is often more useful when the business has a stable instrument but a distinctive customer mix, regulated workflow, high-touch service model, or unusual channel mix. Trend against the same instrument before changing the question. If the instrument changes, mark the break in the series instead of presenting a false continuous trend.
Set operational goals around controllable actions as well as outcomes. For CSAT, review the interaction type and recovery path. For CES, map transfers, repeated authentication, repeated explanations, and channel switching. For NPS, pair the score with the open-ended reason and relationship segment. Bain's own NPS material presents the score as the start of a feedback and improvement system, not a stand-alone dashboard decoration. [1]
Keep benchmark comparisons separate from employee performance rankings when response counts are small or customer mix differs. Small groups can swing on a few responses, and rankings create pressure to suppress invitations or coach customers toward a preferred answer. The safer control is consistent eligibility, transparent sampling, minimum reporting thresholds, and qualitative review of comments.
For a practical implementation companion, see customer satisfaction measurement. For dashboard segmentation and operational reporting context, see customer service metrics dashboards.
A measurement plan for a support team
Start by naming the decision each metric should support. A post-resolution CSAT can inform interaction coaching and recovery review. A task-level CES can identify a broken workflow or unnecessary transfer. A relationship NPS can inform broader product, service, and retention conversations. If one survey asks all three, keep the measures in separate fields and preserve the sequence.
Use a small pilot to test comprehension and routing before setting a target. Check whether customers know which event they are rating, whether the response labels are understandable, and whether the invitation reaches the intended population. Review missing responses by channel and customer segment, not only as one overall rate.
Then publish a measurement note beside the dashboard. It should state the question, scale, formula, dates, population, response handling, and any changes since the previous period. This makes the number auditable and prevents a later reader from mistaking a local operating baseline for an industry norm.
Consolidated statistics table
| Statistic or rule | Value | Source and meaning |
|---|---|---|
| NPS recommendation scale | 0 to 10 | Bain method, from not at all likely to extremely likely [1] |
| NPS promoter band | 9 to 10 | Bain classification [1] |
| NPS passive band | 7 to 8 | Bain classification [1] |
| NPS detractor band | 0 to 6 | Bain classification [1] |
| CSAT example scale | 1 to 5 | Qualtrics example, very unsatisfied to very satisfied [4] |
| CSAT top-box example | 4 and 5 | Qualtrics calculation example [4] |
| CES response direction | Very easy to very difficult | Qualtrics example wording [6] |
These are instrument rules and source examples, not universal performance benchmarks. They should not be presented as targets for every company or channel.
Sources
- Bain, Introducing the Net Promoter System, 2016. NPS question, bands, and calculation.
- Bain, NPS: The next Six Sigma?, historical explanation of the recommendation question and formula.
- Medallia, CSAT: How to Measure and Improve the Customer Service Experience, 2024. CSAT definition and calculation framing.
- Qualtrics, What is CSAT and How Do You Measure It?, 2020. Example 1 to 5 scale, top-box rule, and interaction timing.
- Medallia, Customer Satisfaction Glossary, updated 2026. Distinguishes CSAT and CES aggregation options.
- Qualtrics, Customer Effort Score, 2018. CES definition and example wording.
- Harvard Business Review, Stop Trying to Delight Your Customers, 2010. Historical effort research context.
- Pew Research Center, Writing Survey Questions, survey wording, order, response options, and mode guidance.
- Pew Research Center, Questionnaire Design and Translation, questionnaire design and comparability guidance.
- AAPOR, Response Rates and Survey Quality, response-rate definitions and limits as a quality proxy.
- AAPOR, Standard Definitions, coverage, measurement, and nonresponse error.
- Pew Research Center, Why do some open-ended survey questions result in higher item nonresponse rates?, item nonresponse and question burden.
- Pew Research Center, Survey question wording, wording effects and neutral question construction.
- Keiningham et al., A Longitudinal Examination of Net Promoter and Firm Revenue Growth, peer-reviewed discussion of NPS and growth measurement limits.
Frequently Asked Questions
Is there one good CSAT, NPS, or CES benchmark for every support team?
No. A benchmark is only meaningful when its metric definition, question wording, scale, audience, timing, channel, response handling, and calculation method are comparable. Build a stable internal baseline when those conditions do not match.
Are CSAT, NPS, and CES interchangeable?
No. CSAT measures satisfaction with a defined experience, NPS measures recommendation or advocacy, and CES measures effort in a task or interaction. A dashboard may show all three, but it should not combine them into one unlabeled score.
What should a benchmark report disclose?
At minimum, disclose the exact question, scale, formula, field dates, invitation population, usable responses, response-rate definition, collection mode, exclusions, segments, and known limitations. AAPOR's guidance supports treating response rate as one part of a broader survey-quality assessment.
Should NPS be sent after every support ticket?
Not automatically. A post-ticket survey can measure the customer's view of that interaction, but it may not represent the broader relationship that a relationship NPS question intends to measure. Choose timing based on the decision and label the event clearly.
What is the safest first step when a team has no benchmark?
Choose one clearly worded instrument, pilot it for comprehension and coverage, document the calculation, and establish an internal baseline. Improve the measurement process before comparing the result with an external number.
A practical next step
If your team needs help translating support coverage goals into a measurement and staffing plan, contact CustomerCareStaff to discuss the assumptions and evidence behind your options.