CCDB evaluates debate adjudication at three levels: the final outcome, the local process and individual contribution. The dataset was first released in September 2026 and was not previously public, so models released before that date cannot be contaminated; models released afterwards may be regarded as uncontaminated if they follow standard training-data disclosure practices.
Input format, motion and the complete match transcript. Output p ∈ [0, 1] in steps of 0.01; p > 0.5 predicts the affirmative, p < 0.5 the negative, and an extracted 0.5 is counted as incorrect rather than as a parsing failure, since judges may not tie.
Metrics winner accuracy. Baselines always affirmative, random.
Input the current stage and the whole preceding transcript, with the evaluated side and role identified, plus the judges' rubric. Output an integer score from 1 to 10. The reference is the mean of the three judge scores.
Metrics MSE; Pearson r and Spearman ρ; tendency agreement (all stages and non-tie stages). Baselines random, structure-aware random.
Input as Task 1, in a separate call. Output an integer allocation of nine votes across debaters. The reference is the aggregated nine best-debater votes of the three judges.
Metrics top-1 accuracy (ties in the human maximum count as correct); normalised vote-share MAE; vote-share r and ρ. Baselines random, uniform votes, third speaker on the model-predicted winning side (4v4 subset only). The third speaker on the ground-truth winning side is reported separately as a non-comparable oracle reference.
| Task performance | Battlefield judgment | Degree of advancement | Score |
|---|---|---|---|
| Perfectly completed | Correct and important battlefield | Decisive advancement | 10 |
| Perfectly completed | Correct and important battlefield | Major advancement | 9 |
| Well completed | Correct and important battlefield | Substantial advancement | 8 |
| Well completed | Correct and important battlefield | Effective advancement | 7 |
| Well completed | Correct and important battlefield | An attempt to advance | 6 |
| Well completed | Correct but secondary battlefield | Effective advancement | 6 |
| Approximately completed | Correct but secondary battlefield | An attempt to advance | 5 |
| Approximately completed | Correct but largely irrelevant battlefield | An attempt to advance | 4 |
| Not completed | Incorrect battlefield | An attempt to advance | 3 |
| Not completed | Incorrect battlefield | No advancement | 2 |
| Not completed | Incorrect battlefield | Counterproductive effect | 1 |
| opus-5 | .662 |
| gpt-5.6-sol | .601 |
| gemini-3.5-flash-lite | .595 |
| sonnet-5 | .595 |
| deepseek-v4-flash | .588 |
| haiku-4.5 | .588 |
| gpt-5.6-luna | .561 |
| Always affirmative | .507 |
| Random | .500 |
| Predictor | MSE | r | ρ | Agr. | Non-tie |
|---|---|---|---|---|---|
| deepseek-v4-flash | 2.72 | .250 | .253 | .402 | .638 |
| deepseek-v4-flash (R) | 2.54 | .241 | .241 | .403 | .602 |
| deepseek-v4-pro | 2.13 | .231 | .240 | .434 | .596 |
| gemini-3.5-flash-lite | 1.91 | .184 | .181 | .391 | .529 |
| gpt-5.6-luna | 2.52 | .232 | .237 | .360 | .579 |
| Random | 9.69 | ≈0 | ≈0 | .239 | .451 |
| Structure-aware random | 6.64 | .031 | .050 | .291 | .634 |
| Predictor | Top-1 acc. | MAEvote | Vote-share r | Vote-share ρ | Pred.-winner 3rd speaker |
|---|---|---|---|---|---|
| deepseek-v4-flash | .432 | .125 | .481 | .488 | .457 |
| deepseek-v4-flash (R) | .438 | .129 | .456 | .470 | .461 |
| deepseek-v4-pro | .558 | .118 | .549 | .534 | .491 |
| gemini-3.5-flash-lite | .345 | .136 | .393 | .425 | .474 |
| gpt-5.6-sol | .568 | .117 | .561 | .560 | .439 |
| gpt-5.6-luna | .432 | .122 | .508 | .524 | .422 |
| opus-5 | .520 | .113 | .578 | .591 | .526 |
| sonnet-5 | .419 | .126 | .509 | .515 | .491 |
| haiku-4.5 | .399 | .138 | .395 | .380 | .457 |
| Random | .183 | — | ≈0 | ≈0 | .401 |
| Uniform votes | — | .154 | — | — | — |
| Winning-side third speaker (oracle; 4v4 only) | .569 | .156 | .500 | .435 | .569 |