Quick answer: CSAT measures satisfaction with something that just happened, usually on a 1 to 5 scale, asked right after the moment. NPS asks how likely a customer is to recommend you, on a 0 to 10 scale, and reports the share of promoters minus the share of detractors. Use CSAT for diagnosing specific interactions and NPS for tracking the relationship over time. The real difference is scope rather than quality: neither metric is a fact about your customers, and both carry known distortions worth understanding before you build targets around them.
Most comparisons of these two metrics stop at the formulas. This one spends its middle section on what the academic literature found when somebody checked the famous claim behind NPS, because that changes how much weight the number deserves in a customer experience program.
What Each One Actually Asks
The CSAT question
“How satisfied were you with today’s service?” Answers usually run from very dissatisfied to very satisfied across five points, sometimes seven, occasionally a simple thumbs up or down. CSAT measures satisfaction with the thing the customer just experienced, nothing wider.
The wording ties the answer to a moment. That is the whole design intent, and it is why CSAT gives you something specific to act on rather than a general mood reading.
The NPS question
“How likely are you to recommend us to a friend or colleague?” Zero through ten, with no reference to any particular event. Net Promoter Score is the full name, shortened almost everywhere to the acronym.
Nothing anchors this to a moment, which is deliberate. NPS aims at the whole relationship a customer has with you rather than at one exchange, and that scope choice explains most of its behavior.
How Each Is Calculated
Working out a CSAT satisfaction score
Take the customer responses you count as positive, divide by total responses, multiply by a hundred. Simple, though the definition of positive varies: some organizations count only the top box, others the top two, and comparing across companies gets meaningless quickly for that reason alone. Our guide to calculating it covers the variations.
Working out an NPS figure
Ratings of 9 and 10 are promoters. Ratings of 7 and 8 are passives. Everything from 0 to 6 counts as a detractor. Subtract the detractor percentage from the promoter percentage and you get a figure between negative 100 and positive 100.
Read that bucketing again, because it matters more than anyone tells you when they sell you the method. Our NPS explainer walks through the arithmetic step by step if you want it slower. Three groups, eleven possible answers, and an arithmetic step that discards the distance between responses sitting inside each group.
Side-by-Side
| Aspect | CSAT | NPS |
| What it asks | Satisfaction with one event | Willingness to recommend |
| Scale | Usually 1 to 5 | 0 to 10 |
| Range of result | 0 to 100% | Negative 100 to positive 100 |
| Attached to | A specific interaction | The relationship overall |
| Best timing | Immediately after | Periodic, not event-driven |
| Diagnostic value | High, points at a moment | Low, points at a mood |
| Response rates | Higher, question feels relevant | Lower, feels like an interruption |
| Comparable externally | Poorly, definitions vary | Somewhat, method is standardized |
| Main weakness | Recency and mood effects | Discards most of the scale |
Everything your team needs in one platform
Where Each One Belongs
CSAT after a specific transaction
CSAT is best used within minutes of an interaction ending, while memory is fresh and the customer can still name what happened. Service conversations, deliveries, appointments, onboarding steps. Anywhere a customer has just finished doing something with you and can still recall the detail of how it went.
NPS for the relationship as a whole
NPS is designed for periodic measurement, quarterly or twice a year, sent to a sample rather than to everybody. Asking it after every exchange turns a relationship question into a transaction question and produces noise instead of customer sentiment you can trust.
The Statistical Problems Nobody Mentions
Most of the scale gets thrown away
Here is the objection I find hardest to answer on behalf of NPS. It collects an eleven-point scale and then treats a 6 and a 0 as identical. Somebody mildly lukewarm and somebody actively hostile land in one bucket, contributing the same amount.
Meanwhile 7s and 8s contribute nothing at all. A company could move a third of its customer base from 7 to 8, a real improvement, and see no movement whatsoever in the reported figure.
Try it with numbers. Take a hundred responses averaging 7.4, then imagine every customer shifting up by half a point to 7.9. Customer satisfaction has plainly risen, the average confirms it, and the reported result does not budge because almost nobody crossed a bucket line. Now move eight people from 8 to 9 instead, with the average barely changing, and the figure jumps eight points. NPS responds to boundary crossings rather than to sentiment.
Small samples make the number jump
Because NPS subtracts one percentage from another, a handful of responses near a bucket boundary swings it noticeably. Two people moving from 8 to 9 in a sample of fifty changes the result by four points, which somebody will then explain in a board meeting as the effect of a project.
Who answers is not random
Both metrics suffer here. A customer with strong feelings responds; the indifferent majority does not. That skews results toward extremes in ways no amount of careful survey design fully removes.
It gets worse when response rates fall, because the remaining respondents are increasingly self-selected. A program with a 4% response rate is not measuring your customer base, it is measuring the small group motivated enough to reply, and those two populations are not the same. Track your response rate alongside the result and treat any sharp change in one as a reason to distrust movement in the other.
What the Research Actually Says
The original claim
Frederick Reichheld introduced Net Promoter Score in Harvard Business Review in December 2003, arguing it was the one number that could predict firm growth better than any other. That framing is why the idea spread so fast, and why so many boardrooms treat it as settled.
What replication found
In 2007, Keiningham, Cooil, Andreassen and Aksoy published a longitudinal examination in the Journal of Marketing, using data from 21 firms and more than 15,500 interviews in the Norwegian Customer Satisfaction Barometer. Working in the same industries Reichheld held up as exemplars, they failed to replicate the claimed superiority over other loyalty measures, including plain satisfaction. Their paper won the Marketing Science Institute’s H. Paul Root Award that year, which is not the reception a fringe critique receives.
I would not read that as “NPS is useless.” I read it as: the claim that made it famous did not survive independent checking, so treat the result as one signal among several rather than as the number that governs your customer experience strategy.
Worth being fair about CSAT’s own weakness here too, since it has one. Asked immediately after an interaction, it captures mood as much as merit, and a customer who got bad news politely delivered will often mark you down for the news. That is why a falling CSAT during a price change or a policy tightening usually reflects the change rather than your team, and reading it otherwise leads to coaching people for something outside their control.
There is a reasonable counterargument to the NPS criticism as well, and I will give it fairly. A single standardized question asked the same way everywhere has genuine coordination value inside a large organization, whatever its statistical shortcomings. Everyone knows what it means, nobody argues about definitions, and it travels well between departments. That is worth something real. It is just not the same thing as being the best available predictor of growth, which was the original claim.
Where Effort Sits Alongside Both
A third measure deserves mention, since the CSAT CES pairing often works better than CSAT or NPS alone. Customer Effort Score asks how easy you made it to get something resolved, and the research behind it found effort forecasts future behavior more reliably than delight does.
For a service operation specifically, I would put effort first, CSAT second, and NPS a distant third. Our breakdown of effort scoring covers why.
How to Run Both Without Exhausting People
Sample rather than asking everyone
Survey fatigue is real, and it degrades data quality before it degrades response rates, which is why it goes unnoticed for so long. Send NPS to a rotating fraction of your customer base each quarter. Send CSAT after a share of interactions, not all of them.
Time each one differently
Immediately after the event for CSAT. Well away from any event for NPS, otherwise you are measuring the last thing that happened and calling it loyalty.
There is a related trap in who receives what. Sending the relationship question only to customers who recently contacted your team skews it toward people with a problem, which is the opposite of the representative sample NPS assumes. Draw from your whole base, including the quiet majority who have not needed you lately.
Make the follow-up question earn its place
One open text box asking why. That free-text feedback is worth more than the number attached to it, and reading fifty comments teaches you more than a dashboard ever will. Tools such as SurveyMonkey handle distribution fine; what they cannot do is the reading.
Keep the box optional and short. Making it mandatory raises abandonment noticeably, and a forced comment tends to be a single word rather than the explanation you wanted. Somebody senior should read a sample of these every month, unedited, without a summary layer in between. The habit costs an hour and it changes what leadership believes about the operation more than any chart has ever managed.
Benchmarks, and Why to Distrust Them
Published benchmarks for both metrics are shakier than they look. CSAT figures vary with whether an organization counts one box or two. NPS figures vary with culture: respondents in some countries avoid extreme ratings as a matter of habit, which drags scores down without anything being wrong.
Comparing your own number against last quarter’s is defensible. Comparing it against an industry average from a market research summary is mostly theater, and I would resist doing it in front of an executive team who will remember the figure.
Survey channel matters too, and it rarely appears in the footnotes. The same customer asked by email, by text message, and by an agent at the end of a conversation will often answer differently, with agent-administered questions running noticeably warmer for obvious reasons. If you change how you collect feedback, expect a step change that has nothing to do with your service, and annotate the chart so nobody spends a quarter explaining it.
Frequently Asked Questions
Can you use both at once?
Yes, and most mature operations do. Run satisfaction after individual interactions to find what needs fixing, and the recommendation question quarterly to track whether the relationship is improving. The two answer separate questions, so results diverging is informative rather than contradictory. What you should avoid is asking both in the same survey, which lengthens it, reduces completion, and leaves the customer unsure whether they are rating today’s interaction or the company as a whole.
What counts as a good score?
For satisfaction, most support operations land between 75% and 90% positive, though the figure depends heavily on whether you count the top box or the top two. For NPS, anything above zero means promoters outnumber detractors, and above 50 is genuinely strong. Treat these as rough orientation instead of as targets. Your own trend line matters far more than any published average, since methodology differences between organizations make cross-company comparison unreliable in almost every case.
How many responses do you need for the number to mean anything?
More than most teams collect. For a stable reading you generally want at least a few hundred responses per period, and considerably more if you plan to segment by product or region. Below roughly a hundred, normal variation produces movements that look like trends and are not. Where your sample is small, report a rolling average across several periods rather than a single figure, and resist the urge to explain every wobble to somebody senior who asked.
Does a low score always mean something is wrong?
Not necessarily. Cultural response patterns, question wording, survey channel, and even the time of day all move results independently of anything you did. A drop is a prompt to investigate rather than proof of a problem. Check whether your sample composition changed, whether a survey tool was altered, and whether one specific customer segment drove the whole movement, before concluding that your service quality actually declined. Investigate first, explain second.
Which one should a small operation start with?
Start with satisfaction after interactions. It gives you something you can act on tomorrow, response rates are higher because the question feels relevant, and setup is straightforward. Add the relationship measure later, once you have enough customers for a quarterly sample to be meaningful. Running that second measure across a base of two hundred people produces a number swinging wildly from quarter to quarter, which teaches you very little and wastes goodwill.
Do these measures work for B2B?
They work, with adjustments. Buying groups mean several people at one account hold different views, so an account-level view matters more than an individual average. Response rates tend to be lower and samples smaller, which amplifies the volatility problems described above. Many B2B operations get more value from structured account reviews than from either metric, using the numbers as a supporting signal instead of the headline finding they report upward.
How often should the relationship question be sent?
Quarterly suits most operations, twice yearly for longer purchase cycles. Sending it monthly produces fatigue without producing new information, since attitudes toward a company rarely change that fast. Rotate the sample so nobody receives it repeatedly. If you also run transactional surveys, keep a gap of at least a couple of weeks between the two, so no customer experiences your feedback program as a constant stream of requests they eventually start ignoring.
What should you do with detractors?
Contact them, quickly, and ask what happened. Closed-loop follow-up is the part of the practice with the strongest support behind it, and it produces something the score alone never does: a specific fault you can repair. Aim to reach them within 48 hours while the situation is fresh and the customer still remembers specifics. Track how many were contacted and what changed as a result, because that follow-through is where the value sits.
Is any single number enough to run customer experience on?
No, and treating one that way is the common failure. A score compresses a complicated picture into a digit, which is useful for tracking direction and useless for deciding what to do next. Pair whichever metric you choose with operational data, free-text comments, and conversation analysis. The number will tell you whether something changed. Everything around it is what gives you both a reason and a sensible place to start looking for it.
Want measurement you can actually act on?
Voiso records and transcribes conversations, runs post-interaction surveys automatically, and puts the results next to the operational data that explains them, so a falling score comes with the recordings behind it instead of as a mystery. Talk to the Voiso sales team about what you measure today, and bring your awkward questions about benchmarks.
Sources
- Keiningham, T. L., Cooil, B., Andreassen, T. W. and Aksoy, L. (2007). A Longitudinal Examination of Net Promoter and Firm Revenue Growth. Journal of Marketing 71(3): 39-51: https://journals.sagepub.com/doi/abs/10.1509/jmkg.71.3.039
- Reichheld, F. F. (2003). The One Number You Need to Grow. Harvard Business Review 81 (December): 46-54
- Dixon, M., Toman, N. and DeLisi, R. The Effortless Experience (CEB/Gartner research on effort as a predictor of future behavior)