Documentation · 06 of 07

Chinese Competitive Debating Dataset and Benchmark

Zongrui Yang, Haoyuan Li, Zhongsheng Wang, Zhirui Zeng, Pengqian Han, Yi Zhou, Yuting Wang, Jiamou Liu · Under review as a conference paper at ICLR 2027 · preprint arXiv:2609.21637

Download PDF arXiv 2609.21637 Cite
01 · Abstract

Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models’ understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models’ understanding of interactive argumentation and their agreement with professional judges.

02 · Contributions

Fine-grained, post-verified data

ASR on the recordings, segmentation by stage and clash, and manual verification of all data.

Rubric-based professional scoring

A unified standard agreed with experts first, then more than 180 real matches scored against it, with rationales and votes retained.

Process-level supervision

Supervision refined to the competition-stage level, so models can be trained and evaluated in local debate contexts.

Large-scale expert supervision

About three million Chinese characters of verified transcripts and judgments from more than 120 judges selected under strict criteria.

An LLM benchmark for process understanding

Winner prediction, stage scoring, fine-grained capability evaluation and human–model agreement, evaluated across multiple models.

03 · Comparison with existing Chinese debate datasets
DatasetCorpus sizeCorpus granularitySupervision sourceSupervision granularitySupervision scale
ORCHID (2023)1,218 matchesUtterance; no stage segmentationOfficial outcomes; no manual labelsMatch level1,218 matches
CDWC (Chen et al., 2026a)94 matchesUtterance; no stage segmentationGPT-4.5-Preview scores; 3 on-site judges for outcomes; non-expert validationUtterance-level support / response / language; match outcome94 match-level; 216 utterance-level
DEFINED (2026)108 matchesStage / utterance3–5 raters per statement; 10 graduate students for fine-grained relabellingStatement level706 coarse; 120 fine-grained
CEDAR (2026)600 matchesSentence / utterance; no stage segmentation10 trained annotators in two groups with senior adjudicationSentence-level claim / stance / evidence / rhetoric; match result8,251 claims; 5,173 evidence; 1,126 rhetoric; 3,631 argument pairs
Conch (2026b)3 matchesSession → turn → blockLLM extraction validated by 3 expertsBlock / clash / path structureFull analysis of 3 matches
CCDD (ours)182 matches, organised competitionsMatch → stage → paired clash / QA unit120 on-site professional judges, 3 per match, predefined rubricMatch + stage + paired stage + speaker role148 winner; 2,693 stage; 1,345 paired-stage; 1,044 best-debater labels