Documentation · 03 of 07

Data collection

Rather than relabelling public matches after the fact, we first agreed a unified rubric with senior debate experts, then organised real competitions under it. Judges scored every stage during live adjudication, and their votes, ballots and rationales were retained.

01 · The competitions
182
Matches organised: 150 4v4 and 32 2v2 cup matches
3
Judges per match, drawn from a pool of 120+
168
Complete adjudication instances retained: 148 ordinary, 20 appeal
3.85M
Chinese characters: 2.60M transcript, 1.25M judge commentary

These were real competitions whose outcomes mattered to participants independently of this research. Each judge independently scored every stage against the rubric and cast an impression vote, a stage vote and a deciding vote, then delivered a post-match commentary averaging about 3,000 Chinese characters. Score sheets were submitted before rationales were given orally, so one judge's opinion could not shift another's scores.

02 · Match format

Constructive phase

Six segments: affirmative constructive, negative questioning, negative constructive, affirmative questioning, and a questioning summary from each side. Establishes definitions, the criterion and the main arguments.

Clash phase

Five segments: head-to-head exchange, affirmative and negative cross-examination, and a cross-examination summary from each side. Attack, defence and comparison on the disputes already established.

Concluding phase

Free debate on separate team clocks, then a closing speech from each side that reviews the major disputes and gives judges a final comparative framework.

03 · Processing pipeline
StepOperationNotes
01Recording and formsOne MP4 per match with all audio and the timer feed; three XLSX voting forms with stage scores, vote placements and best-debater nominees.
02FilteringMatches with incomplete judge forms removed; 168 complete adjudication instances retained.
03ASR and constrained repairTencent Meeting ASR transcripts (about 3.85 million characters from 12,120 minutes of video), corrected by Claude Fable 5 under a constrained prompt while keeping every source-line index for sentence-level auditing.
04Stage segmentationChair commands and the timing track split the sentence sequence into eight stage types with name, type, acting side and utterance list.
05Score alignmentStage rows in each of the three score sheets aligned one-to-one with transcript stages; questioning and answering sides scored separately in interactive stages.
06Manual verificationAll text entries of the retained matches were reviewed with a dedicated review tool for transcription accuracy and stage boundaries; 71 stage segmentations were adjusted and 304 transcription errors corrected. Retained source-line numbers trace every sentence back to the original ASR record.
07Clash segmentationQuestioning and cross-examination split into question–answer turns; free debate split into attack–defence rounds.
04 · Appeals and label reliability

Appeal mechanism

Debaters who disagree with a result may ask three judges with stronger credentials to re-adjudicate the same recording independently, replacing the original result. Appeals are deliberately easy to initiate. Of 182 matches, 20 were appealed and 7 results overturned, so 96.2% of original results stood.

Debate has no ground-truth winner. Votes record which side persuaded more judges; a dissenting judge is not thereby wrong.

Unanimous 3:0 decisions55.4%
Split 2:1 decisions44.6%
Pairwise judge agreement on winner70.3%
Estimated single-judge / panel accuracy81.9% / 91.3%
MSE among the three judges' stage scores0.759
Same score band, all three / two / none52.8 / 44.3 / 2.8%
ICC(1,3), raw scores / tendencies0.523 / 0.630
Every judge gave ≥1 vote to the eventual best debater87.8%