Research question
How should a customer service team select cases for quality review without presenting a convenient subset as representative of every customer? Sampling affects what a team sees, what it misses, and how confidently it can act. This report covers support contacts, case reviews, and customer feedback. It does not prescribe one sample size or claim a universal confidence threshold.
Start with the population
Name the population before selecting records. It might be all eligible cases closed during a period, all chats offered a survey, or all escalations received by a specialist queue. State whether duplicates, spam, abandoned contacts, and cases still open are included. A population defined after viewing results can quietly exclude inconvenient records.
The American Association for Public Opinion Research distinguishes several response and cooperation concepts and emphasizes transparent disposition reporting. The same discipline helps support reviews. For a case audit, document created, eligible, selected, available, and reviewed counts. For feedback, document invitations, deliveries, responses, and exclusions. A denominator without its path is not enough.
Match method to decision
Use a probability-oriented sample when the question is about an estimated property of a defined population and selection probabilities can be described. Use a purposeful sample when the question is to find rare failures, inspect a new workflow, or learn about a known risk. Neither is automatically superior.
| Goal | Suitable approach | Main statement it supports |
|---|---|---|
| Estimate a population property | Random or stratified selection | An estimate for the defined population, with uncertainty |
| Find rare high-risk cases | Oversample a risk stratum | Findings about the reviewed risk group |
| Understand a new workflow | Purposeful case review | Themes and defects in tested scenarios |
| Monitor change over time | Stable repeated design | Direction under comparable methods |
If strata are oversampled, weight or report the strata separately when estimating the population. Do not average reviewed groups as though each had the same population share.
Methodology
Write the selection rule in advance. Include the randomization source or queue ordering, the period, the unit, replacement rule, and treatment of unavailable records. If a reviewer chooses cases from a dashboard, the review is not random unless the dashboard selection process is documented.
Use a sample frame that covers the channels and case types in scope. If phone transcripts are unavailable, do not call the review “all support.” State the coverage gap. For customer surveys, report the invitation method and field period. AAPOR warns that response rate alone does not establish absence of nonresponse bias. The same applies to a high case-review completion rate when cases are missing systematically.
Define reviewer agreement and adjudication before scoring. Two reviewers can disagree because the rubric is unclear, not because one is careless. Preserve the first independent scores and the reason for adjudication. This makes changes to the rubric visible.
Analyze without false precision
Report counts with percentages and the denominator. Avoid more decimal places than the sample supports. Small slices can change substantially when one record moves. When a measure is an estimate, provide an uncertainty statement appropriate to the design. Do not use a generic margin of error for a purposeful sample or a clustered sample without checking the assumptions.
Separate descriptive findings from causal claims. If low documentation and repeat contact appear together, that is a useful association to investigate. It does not prove that note quality caused the repeat. Check case complexity, staffing, policy changes, and channel differences.
Protect customers in review data. The NIST Privacy Framework recommends identifying privacy risk in collection, processing, and sharing. Redact unnecessary personal information, limit access, and set a retention rule. If a quality reviewer needs the text of a message, another analyst may need only a coded field.
Set a stopping and replacement rule before collection. If a selected record is unavailable, do not quietly replace it with the next convenient record. Record why it was unavailable and apply the predefined rule. If replacements are used, explain how they can change the sample. For longitudinal reviews, preserve the method even when the queue changes so that a trend break is visible.
Quality review should also protect the people whose work is being reviewed. Give reviewers a rubric that separates observable behavior from judgment, and let agents challenge an incorrect factual premise without changing the original score. This reduces measurement error and prevents a review process from rewarding documentation theater.
Limitations and conclusion
Every sample has a frame, scope, and missingness pattern. A carefully selected sample cannot repair a population that excludes an entire channel. A large feedback count can still be biased by who chooses to respond. A purposeful review can find a severe defect without estimating its frequency.
A defensible support sampling plan names the population, matches selection to the decision, records exclusions and missingness, preserves reviewer independence, and limits the claim to the evidence. The result is more useful when uncertainty is visible than when a precise-looking number obscures how the records were chosen.
Sources
- AAPOR, Standard Definitions, response outcomes and transparent disposition reporting.
- AAPOR, Response Rates, limits of response rate as a quality measure.
- NIST Privacy Framework, privacy risk in data collection and use.
- Pew Research Center, Writing Survey Questions, survey design and wording effects.
- US Census Bureau, Research and Methodology, quality and methodological transparency context.
Is a larger support sample always better?
No. Coverage, selection, measurement quality, and missingness matter. A large biased sample can answer the wrong question precisely.