Zongrui Yang, Haoyuan Li, Zhongsheng Wang, Zhirui Zeng, Pengqian Han, Yi Zhou, Yuting Wang, Jiamou Liu · Under review as a conference paper at ICLR 2027 · preprint arXiv:2609.21637
Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models’ understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models’ understanding of interactive argumentation and their agreement with professional judges.
ASR on the recordings, segmentation by stage and clash, and manual verification of all data.
A unified standard agreed with experts first, then more than 180 real matches scored against it, with rationales and votes retained.
Supervision refined to the competition-stage level, so models can be trained and evaluated in local debate contexts.
About three million Chinese characters of verified transcripts and judgments from more than 120 judges selected under strict criteria.
Winner prediction, stage scoring, fine-grained capability evaluation and human–model agreement, evaluated across multiple models.
| Dataset | Corpus size | Corpus granularity | Supervision source | Supervision granularity | Supervision scale |
|---|---|---|---|---|---|
| ORCHID (2023) | 1,218 matches | Utterance; no stage segmentation | Official outcomes; no manual labels | Match level | 1,218 matches |
| CDWC (Chen et al., 2026a) | 94 matches | Utterance; no stage segmentation | GPT-4.5-Preview scores; 3 on-site judges for outcomes; non-expert validation | Utterance-level support / response / language; match outcome | 94 match-level; 216 utterance-level |
| DEFINED (2026) | 108 matches | Stage / utterance | 3–5 raters per statement; 10 graduate students for fine-grained relabelling | Statement level | 706 coarse; 120 fine-grained |
| CEDAR (2026) | 600 matches | Sentence / utterance; no stage segmentation | 10 trained annotators in two groups with senior adjudication | Sentence-level claim / stance / evidence / rhetoric; match result | 8,251 claims; 5,173 evidence; 1,126 rhetoric; 3,631 argument pairs |
| Conch (2026b) | 3 matches | Session → turn → block | LLM extraction validated by 3 experts | Block / clash / path structure | Full analysis of 3 matches |
| CCDD (ours) | 182 matches, organised competitions | Match → stage → paired clash / QA unit | 120 on-site professional judges, 3 per match, predefined rubric | Match + stage + paired stage + speaker role | 148 winner; 2,693 stage; 1,345 paired-stage; 1,044 best-debater labels |