Chinese Competitive Debating Dataset & Benchmark

Together with debate tournament organisers and debate clubs we built the Chinese Competitive Debating Dataset (CCDD) and Benchmark (CCDB). We organised 182 competitive debates; after removing matches with missing records, 148 independent matches and an appeal set of 20 re-adjudicated matches remain. Every match was scored by three professional judges. For each judge we release stage-by-stage scores, totals, votes, appeal records, best-debater ballots and an adjudication rationale averaging 3,000 Chinese characters, together with aggregated summaries.

Read the paper Data documentation License & terms
CCDD — corpus and supervision at a glance v1.0 Sep 2026 Sheet 01
01 · Matches
148
Independent matches, plus 20 appeal re-adjudications
02 · Stages
2,698
Stage-level instances, each with three rubric scores
03 · Clashes
20,542
Clash-level exchanges, segmented from interactive stages
04 · Judges
120+
Professional judges, three per match, one shared rubric
No.Per-judge recordCoverageRemark
05Stage scores and totals9,163 signalsOne stage text with context, one score from one judge
06Votes, appeals and outcome504 signalsImpression, stage and deciding votes; nine-vote ballot pattern
07Best-debater ballots1,044 labelsNine votes per match, aggregated to a vote distribution
08Adjudication rationale≈3,000 charsPer judge per match; 1.25M characters in total

Clash-level text carries no separate labels; every match and stage carries three or more independent judgments (more where an appeal panel re-scored), so supervised pairs can be composed at multiple granularities.

02 · Why this dataset

Finer corpus granularity

Processing goes down to the individual clash. Earlier debate datasets stop at the stage, yet a single questioning stage may hold dozens of questions. Every speech segment here is split within its stage into question–answer and attack–defence units.

Finer supervision

Most datasets provide only the match result. Our judges scored both sides at every stage during the live adjudication, a process-level supervision signal no other debate dataset offers.

Professional supervision

Some datasets derive labels by post-processing transcripts or by prompting a model. Every label here is an original judgment made by a qualified judge, on site, against one predefined rubric.

03 · Model performance · zero-shot, 148 matches

Task 1 · Winner prediction

Accuracy · axis 0–1
opus-50.662
gpt-5.6-sol0.601
gemini-3.5-flash-lite0.595
sonnet-50.595
deepseek-v4-flash0.588
haiku-4.50.588
gpt-5.6-luna0.561
Always affirmative0.507
Random0.500

Task 2 · Stage scoring

Pearson r with mean judge score · axis 0–0.5
deepseek-v4-flash0.250
deepseek-v4-flash (R)0.241
gpt-5.6-luna0.232
deepseek-v4-pro0.231
gemini-3.5-flash-lite0.184
Structure-aware random0.031
Random0.000

Task 3 · Best debater

Top-1 accuracy · axis 0–1
gpt-5.6-sol0.568
deepseek-v4-pro0.558
opus-50.520
deepseek-v4-flash (R)0.438
deepseek-v4-flash0.432
gpt-5.6-luna0.432
sonnet-50.419
haiku-4.50.399
gemini-3.5-flash-lite0.345
Winning-side 3rd speaker (oracle)0.569
Random0.183

Dashed outlines are structural baselines. The best winner accuracy is 66.2%, the highest stage-score correlation with judges is about 0.25, and the best top-1 best-debater accuracy is 56.8%. Full tables and metric definitions are in the evaluation protocol.

04 · Documentation
01
Data documentation
Files, levels, fields and counts
02
License & terms of use
What you may do with the data
03
Data collection
Competitions, pipeline and label reliability
04
Expert qualifications
Who the judges are and how they were assigned
05
Evaluation protocol
Tasks, metrics, splits and the rubric
06
Paper
Abstract, comparison with prior work, PDF
07
Privacy policy
Debaters, judges and recordings
05 · Citation
arXiv:2609.21637Cite the paper for the dataset and the benchmark; the second entry identifies the dataset release itself.
Paper · CCDD & CCDB
@misc{yang2026ccdd,
  title  = {Chinese Competitive Debating Dataset and Benchmark},
  author = {Yang, Zongrui and Li, Haoyuan and Wang, Zhongsheng and Zeng, Zhirui and Han, Pengqian and Zhou, Yi and Wang, Yuting and Liu, Jiamou},
  year   = {2026},
  eprint = {2609.21637},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url    = {https://arxiv.org/abs/2609.21637}
}
Dataset release · CCDD v1.0
@misc{ccdd2026dataset,
  title  = {CCDD: Chinese Competitive Debating Dataset, version 1.0},
  author = {Yang, Zongrui and Li, Haoyuan and Wang, Zhongsheng and Zeng, Zhirui and Han, Pengqian and Zhou, Yi and Wang, Yuting and Liu, Jiamou},
  year   = {2026},
  publisher = {Liu AI Lab, University of Auckland},
  url    = {https://ccdd.debate.download},
  note   = {Released under CC BY-NC-SA 4.0}
}
06 · Acknowledgements

We thank all friends in the debating community who offered advice, feedback and help with data processing for this benchmark.

We thank Mr. Xi Rui (席瑞), Mr. Wan Huaming (万华明), Ms. Wu Xinning (仵心凝) and other senior figures for their insight and ideas; and Beiming Debate Club (北冥辩论俱乐部), Adventurers' Tavern Debate Club (冒险者酒馆辩论俱乐部), the New Zealand Chinese Debating Association (新西兰华语辩论协会) and other organisations for their support and advice throughout data review, processing, tournament organisation and evaluation design.

Above all, we thank every expert judge who took part. Your long-accumulated debating practice and adjudication experience, and the care you brought to judging each match, form the most important professional foundation of this dataset and benchmark, and its most valuable asset.