Rather than relabelling public matches after the fact, we first agreed a unified rubric with senior debate experts, then organised real competitions under it. Judges scored every stage during live adjudication, and their votes, ballots and rationales were retained.
These were real competitions whose outcomes mattered to participants independently of this research. Each judge independently scored every stage against the rubric and cast an impression vote, a stage vote and a deciding vote, then delivered a post-match commentary averaging about 3,000 Chinese characters. Score sheets were submitted before rationales were given orally, so one judge's opinion could not shift another's scores.
Six segments: affirmative constructive, negative questioning, negative constructive, affirmative questioning, and a questioning summary from each side. Establishes definitions, the criterion and the main arguments.
Five segments: head-to-head exchange, affirmative and negative cross-examination, and a cross-examination summary from each side. Attack, defence and comparison on the disputes already established.
Free debate on separate team clocks, then a closing speech from each side that reviews the major disputes and gives judges a final comparative framework.
| Step | Operation | Notes |
|---|---|---|
| 01 | Recording and forms | One MP4 per match with all audio and the timer feed; three XLSX voting forms with stage scores, vote placements and best-debater nominees. |
| 02 | Filtering | Matches with incomplete judge forms removed; 168 complete adjudication instances retained. |
| 03 | ASR and constrained repair | Tencent Meeting ASR transcripts (about 3.85 million characters from 12,120 minutes of video), corrected by Claude Fable 5 under a constrained prompt while keeping every source-line index for sentence-level auditing. |
| 04 | Stage segmentation | Chair commands and the timing track split the sentence sequence into eight stage types with name, type, acting side and utterance list. |
| 05 | Score alignment | Stage rows in each of the three score sheets aligned one-to-one with transcript stages; questioning and answering sides scored separately in interactive stages. |
| 06 | Manual verification | All text entries of the retained matches were reviewed with a dedicated review tool for transcription accuracy and stage boundaries; 71 stage segmentations were adjusted and 304 transcription errors corrected. Retained source-line numbers trace every sentence back to the original ASR record. |
| 07 | Clash segmentation | Questioning and cross-examination split into question–answer turns; free debate split into attack–defence rounds. |
Debaters who disagree with a result may ask three judges with stronger credentials to re-adjudicate the same recording independently, replacing the original result. Appeals are deliberately easy to initiate. Of 182 matches, 20 were appealed and 7 results overturned, so 96.2% of original results stood.
Debate has no ground-truth winner. Votes record which side persuaded more judges; a dissenting judge is not thereby wrong.
| Unanimous 3:0 decisions | 55.4% |
| Split 2:1 decisions | 44.6% |
| Pairwise judge agreement on winner | 70.3% |
| Estimated single-judge / panel accuracy | 81.9% / 91.3% |
| MSE among the three judges' stage scores | 0.759 |
| Same score band, all three / two / none | 52.8 / 44.3 / 2.8% |
| ICC(1,3), raw scores / tendencies | 0.523 / 0.630 |
| Every judge gave ≥1 vote to the eventual best debater | 87.8% |