Every label in CCDD is an original judgment by a professional debate judge, made on site under one rubric. The judge pool of more than 120 people was selected under strict, uniform criteria.
Each match was assigned three judges from the pool. They scored every stage independently against the rubric, then cast an impression vote, a stage vote and a deciding vote, and nominated best debaters.
Judges submitted their score sheets before explaining their rationale orally, so no judge's opinion could influence the scores of the others. Each rationale averages about 3,000 Chinese characters.
Appealed matches were re-adjudicated from the recording by three judges with stronger credentials than the original panel. Their results replace the original and are marked with a post-appeal identifier.
Before any competition was held we consulted a number of well-known debate experts and agreed a unified stage-level rubric. Judges assess each stage along three dimensions, task performance, battlefield judgment and degree of advancement, and assign a score from 1 to 10. The scores fall into four bands: not completed (1–3), approximately completed (4–5), well completed (6–8) and perfectly completed (9–10).
The full rubric table is reproduced in the evaluation protocol; models are shown the same rubric the judges used.