Documentation · 05 of 07

Evaluation protocol

CCDB evaluates debate adjudication at three levels: the final outcome, the local process and individual contribution. The dataset was first released in September 2026 and was not previously public, so models released before that date cannot be contaminated; models released afterwards may be regarded as uncontaminated if they follow standard training-data disclosure practices.

01 · Tasks
Task 1

Winner-tendency prediction

Input format, motion and the complete match transcript. Output p ∈ [0, 1] in steps of 0.01; p > 0.5 predicts the affirmative, p < 0.5 the negative, and an extracted 0.5 is counted as incorrect rather than as a parsing failure, since judges may not tie.

Metrics winner accuracy. Baselines always affirmative, random.

Task 2

Stage-score prediction

Input the current stage and the whole preceding transcript, with the evaluated side and role identified, plus the judges' rubric. Output an integer score from 1 to 10. The reference is the mean of the three judge scores.

Metrics MSE; Pearson r and Spearman ρ; tendency agreement (all stages and non-tie stages). Baselines random, structure-aware random.

Task 3

Best-debater prediction

Input as Task 1, in a separate call. Output an integer allocation of nine votes across debaters. The reference is the aggregated nine best-debater votes of the three judges.

Metrics top-1 accuracy (ties in the human maximum count as correct); normalised vote-share MAE; vote-share r and ρ. Baselines random, uniform votes, third speaker on the model-predicted winning side (4v4 subset only). The third speaker on the ground-truth winning side is reported separately as a non-comparable oracle reference.

02 · Zero-shot evaluation
03 · Fine-tuning splits
04 · Stage-level rubric used by judges and models
Task performanceBattlefield judgmentDegree of advancementScore
Perfectly completedCorrect and important battlefieldDecisive advancement10
Perfectly completedCorrect and important battlefieldMajor advancement9
Well completedCorrect and important battlefieldSubstantial advancement8
Well completedCorrect and important battlefieldEffective advancement7
Well completedCorrect and important battlefieldAn attempt to advance6
Well completedCorrect but secondary battlefieldEffective advancement6
Approximately completedCorrect but secondary battlefieldAn attempt to advance5
Approximately completedCorrect but largely irrelevant battlefieldAn attempt to advance4
Not completedIncorrect battlefieldAn attempt to advance3
Not completedIncorrect battlefieldNo advancement2
Not completedIncorrect battlefieldCounterproductive effect1
05 · Zero-shot results · 148 matches, 2,693 stage instances
Task 1 · Winner accuracy
opus-5.662
gpt-5.6-sol.601
gemini-3.5-flash-lite.595
sonnet-5.595
deepseek-v4-flash.588
haiku-4.5.588
gpt-5.6-luna.561
Always affirmative.507
Random.500
Task 2 · Stage scoring
PredictorMSErρAgr.Non-tie
deepseek-v4-flash2.72.250.253.402.638
deepseek-v4-flash (R)2.54.241.241.403.602
deepseek-v4-pro2.13.231.240.434.596
gemini-3.5-flash-lite1.91.184.181.391.529
gpt-5.6-luna2.52.232.237.360.579
Random9.69≈0≈0.239.451
Structure-aware random6.64.031.050.291.634
Task 3 · Best debater and vote distribution
PredictorTop-1 acc.MAEvoteVote-share rVote-share ρPred.-winner 3rd speaker
deepseek-v4-flash.432.125.481.488.457
deepseek-v4-flash (R).438.129.456.470.461
deepseek-v4-pro.558.118.549.534.491
gemini-3.5-flash-lite.345.136.393.425.474
gpt-5.6-sol.568.117.561.560.439
gpt-5.6-luna.432.122.508.524.422
opus-5.520.113.578.591.526
sonnet-5.419.126.509.515.491
haiku-4.5.399.138.395.380.457
Random.183≈0≈0.401
Uniform votes.154
Winning-side third speaker (oracle; 4v4 only).569.156.500.435.569