Research question and scope
Published September 9, 2026.
This study asks whether trained reviewers assign the same disposition code to the same closed support case. It covers a named taxonomy, case population, and review window. Agreement does not establish that the taxonomy captures every useful business outcome.
Methodology
Freeze the codebook version and select a stratified sample across queues, issue families, channels, and existing codes. Remove the original code before independent review. Give reviewers the same permitted evidence and collect one primary code plus an uncertainty flag. Adjudicate disagreements only after preserving the independent results.
Measures and analysis
Report raw agreement, category-specific confusion, prevalence, missing evidence, and an appropriate chance-adjusted statistic with its assumptions. Show results for rare categories separately because an overall percentage can hide poor performance. Document codebook changes prompted by the review and retest a fresh sample.
Limitations and inference limits
Reviewers may share training or organizational assumptions. Redacted cases may remove cues used in live work. Chance-adjusted measures are sensitive to category prevalence and should not be read as a universal quality grade. Findings apply only to the tested codebook and evidence set.