Documentation · 02 of 07

License & terms of use

Version 1.0Effective 21 September 2026. Downloading or using the dataset or benchmark code constitutes acceptance of the terms below.
01 · Licenses
Dataset · CCDD

CC BY-NC-SA 4.0

Transcripts, stage segmentation, scores, votes, ballots and rationales. Non-commercial research use with attribution; derivatives under the same terms.

Benchmark code · CCDB

MIT

Prompts, parsing, metrics and split definitions. Free to use, modify and redistribute with the copyright notice.

Raw recordings

Not redistributed

Match recordings contain the voices of debaters and chairs and are held by the authors. Access, if any, is arranged case by case under a data-use agreement.

02 · Terms of use
  1. Attribution. Cite the dataset and benchmark as given on the citation section in any publication, model card or derived resource.
  2. No re-identification. Do not attempt to identify individual debaters, judges or teams from transcripts, scores, ballots or rationales, or to link records to people outside the dataset.
  3. Judgments are not ground truth. Votes and scores record which side persuaded professional judges under a shared rubric. Do not present them as objective correctness of arguments or as assessments of the people involved.
  4. Evaluation integrity. Models evaluated on CCDB should not have been trained on CCDD test folds. Report the split, the parsing rule and the parsing success rate as described in the evaluation protocol.
  5. No training of commercial models. The dataset may not be used, in whole or in part, to train, fine-tune or align models that are offered commercially, whether as a product, an API or a service. This restates the non-commercial condition of the license in the terms most relevant to language models.
  6. No training on the benchmark. Do not train, fine-tune or tune prompts on CCDB test material. A model that has seen the evaluation data no longer measures what the benchmark is designed to measure, and results obtained that way must not be reported as CCDB results. If contamination is discovered after the fact, disclose it alongside any reported numbers.
  7. Derived data. Fine-tuned models, filtered subsets and additional annotations may be shared under the same terms; include a pointer to this page.
  8. Takedown. Participants may request correction or removal of records that concern them. See the privacy policy.