Together with debate tournament organisers and debate clubs we built the Chinese Competitive Debating Dataset (CCDD) and Benchmark (CCDB). We organised 182 competitive debates; after removing matches with missing records, 148 independent matches and an appeal set of 20 re-adjudicated matches remain. Every match was scored by three professional judges. For each judge we release stage-by-stage scores, totals, votes, appeal records, best-debater ballots and an adjudication rationale averaging 3,000 Chinese characters, together with aggregated summaries.
| No. | Per-judge record | Coverage | Remark |
|---|---|---|---|
| 05 | Stage scores and totals | 9,163 signals | One stage text with context, one score from one judge |
| 06 | Votes, appeals and outcome | 504 signals | Impression, stage and deciding votes; nine-vote ballot pattern |
| 07 | Best-debater ballots | 1,044 labels | Nine votes per match, aggregated to a vote distribution |
| 08 | Adjudication rationale | ≈3,000 chars | Per judge per match; 1.25M characters in total |
Clash-level text carries no separate labels; every match and stage carries three or more independent judgments (more where an appeal panel re-scored), so supervised pairs can be composed at multiple granularities.
Processing goes down to the individual clash. Earlier debate datasets stop at the stage, yet a single questioning stage may hold dozens of questions. Every speech segment here is split within its stage into question–answer and attack–defence units.
Most datasets provide only the match result. Our judges scored both sides at every stage during the live adjudication, a process-level supervision signal no other debate dataset offers.
Some datasets derive labels by post-processing transcripts or by prompting a model. Every label here is an original judgment made by a qualified judge, on site, against one predefined rubric.
Dashed outlines are structural baselines. The best winner accuracy is 66.2%, the highest stage-score correlation with judges is about 0.25, and the best top-1 best-debater accuracy is 56.8%. Full tables and metric definitions are in the evaluation protocol.
@misc{yang2026ccdd,
title = {Chinese Competitive Debating Dataset and Benchmark},
author = {Yang, Zongrui and Li, Haoyuan and Wang, Zhongsheng and Zeng, Zhirui and Han, Pengqian and Zhou, Yi and Wang, Yuting and Liu, Jiamou},
year = {2026},
eprint = {2609.21637},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.21637}
}@misc{ccdd2026dataset,
title = {CCDD: Chinese Competitive Debating Dataset, version 1.0},
author = {Yang, Zongrui and Li, Haoyuan and Wang, Zhongsheng and Zeng, Zhirui and Han, Pengqian and Zhou, Yi and Wang, Yuting and Liu, Jiamou},
year = {2026},
publisher = {Liu AI Lab, University of Auckland},
url = {https://ccdd.debate.download},
note = {Released under CC BY-NC-SA 4.0}
}We thank all friends in the debating community who offered advice, feedback and help with data processing for this benchmark.
We thank Mr. Xi Rui (席瑞), Mr. Wan Huaming (万华明), Ms. Wu Xinning (仵心凝) and other senior figures for their insight and ideas; and Beiming Debate Club (北冥辩论俱乐部), Adventurers' Tavern Debate Club (冒险者酒馆辩论俱乐部), the New Zealand Chinese Debating Association (新西兰华语辩论协会) and other organisations for their support and advice throughout data review, processing, tournament organisation and evaluation design.
Above all, we thank every expert judge who took part. Your long-accumulated debating practice and adjudication experience, and the care you brought to judging each match, form the most important professional foundation of this dataset and benchmark, and its most valuable asset.